We consider Pareto front identification for linear bandits (PFILin) where the goal is to identify a set of arms whose reward vectors are not dominated by any of the others when the mean reward vector is a linear function of the context. PFILin includes the best arm identification problem and multi-objective active learning as special cases. The sample complexity of our proposed algorithm is $\tilde{O}(d/\Delta^2)$, where $d$ is the dimension of contexts and $\Delta$ is a measure of problem complexity. Our sample complexity is optimal up to a logarithmic factor. A novel feature of our algorithm is that it uses the contexts of all actions. In addition to efficiently identifying the Pareto front, our algorithm also guarantees $\tilde{O}(\sqrt{d/t})$ bound for instantaneous Pareto regret when the number of samples is larger than $\Omega(d\log dL)$ for $L$ dimensional vector rewards. By using the contexts of all arms, our proposed algorithm simultaneously provides efficient Pareto front identification and regret minimization. Numerical experiments demonstrate that the proposed algorithm successfully identifies the Pareto front while minimizing the regret.
翻译:我们考虑线性赌博机中的帕累托前沿识别问题(PFILin),其目标是在平均奖励向量为上下文的线性函数时,识别出奖励向量不被任何其他臂支配的臂集合。PFILin将最佳臂识别问题和多目标主动学习作为特例包含在内。我们提出算法的样本复杂度为$\tilde{O}(d/\Delta^2)$,其中$d$是上下文的维度,$\Delta$是问题复杂度的度量。该样本复杂度在忽略对数因子的情况下达到最优。我们算法的一个新颖特征在于使用了所有动作的上下文信息。除了高效识别帕累托前沿外,当样本数量大于$\Omega(d\log dL)$($L$为向量奖励的维度)时,我们的算法还能保证瞬时帕累托遗憾的$\tilde{O}(\sqrt{d/t})$界。通过利用所有臂的上下文,我们提出的算法同时实现了高效的帕累托前沿识别和遗憾最小化。数值实验表明,所提算法能在最小化遗憾的同时成功识别帕累托前沿。