We introduce vector optimization problems with stochastic bandit feedback, in which preferences among designs are encoded by a polyhedral ordering cone $C$. Our setup generalizes the best arm identification problem to vector-valued rewards by extending the concept of Pareto set beyond multi-objective optimization. We characterize the sample complexity of ($\epsilon,\delta$)-PAC Pareto set identification by defining a new cone-dependent notion of complexity, called the ordering complexity. In particular, we provide gap-dependent and worst-case lower bounds on the sample complexity and show that, in the worst-case, the sample complexity scales with the square of ordering complexity. Furthermore, we investigate the sample complexity of the na\"ive elimination algorithm and prove that it nearly matches the worst-case sample complexity. Finally, we run experiments to verify our theoretical results and illustrate how $C$ and sampling budget affect the Pareto set, the returned ($\epsilon,\delta$)-PAC Pareto set, and the success of identification.
翻译:我们引入了带随机赌博反馈的向量优化问题,其中设计间的偏好由多面体序锥$C$编码。我们的框架通过将帕累托集概念扩展到多目标优化之外,将最优臂识别问题推广至向量值奖励。通过定义一种新的依赖于锥的复杂度概念——序复杂度,我们刻画了($\epsilon,\delta$)-PAC帕累托集识别的样本复杂度。具体地,我们给出了样本复杂度的间隙依赖下界和最坏情况下界,并证明在最坏情况下样本复杂度随序复杂度的平方增长。此外,我们研究了朴素消除算法的样本复杂度,并证明其几乎匹配最坏情况样本复杂度。最后,我们通过实验验证理论结果,并说明$C$与采样预算如何影响帕累托集、返回的($\epsilon,\delta$)-PAC帕累托集以及识别成功率。