This paper considers a stochastic Multi-Armed Bandit (MAB) problem with dual objectives: (i) quick identification and commitment to the optimal arm, and (ii) reward maximization throughout a sequence of $T$ consecutive rounds. Though each objective has been individually well-studied, i.e., best arm identification for (i) and regret minimization for (ii), the simultaneous realization of both objectives remains an open problem, despite its practical importance. This paper introduces \emph{Regret Optimal Best Arm Identification} (ROBAI) which aims to achieve these dual objectives. To solve ROBAI with both pre-determined stopping time and adaptive stopping time requirements, we present an algorithm called EOCP and its variants respectively, which not only achieve asymptotic optimal regret in both Gaussian and general bandits, but also commit to the optimal arm in $\mathcal{O}(\log T)$ rounds with pre-determined stopping time and $\mathcal{O}(\log^2 T)$ rounds with adaptive stopping time. We further characterize lower bounds on the commitment time (equivalent to the sample complexity) of ROBAI, showing that EOCP and its variants are sample optimal with pre-determined stopping time, and almost sample optimal with adaptive stopping time. Numerical results confirm our theoretical analysis and reveal an interesting "over-exploration" phenomenon carried by classic UCB algorithms, such that EOCP has smaller regret even though it stops exploration much earlier than UCB, i.e., $\mathcal{O}(\log T)$ versus $\mathcal{O}(T)$, which suggests over-exploration is unnecessary and potentially harmful to system performance.
翻译:本文研究一个具有双重目标的随机多臂老虎机问题:(i) 快速识别并锁定最优臂,(ii) 在连续 $T$ 轮中最大化累积奖励。尽管每个目标均已得到充分研究——(i) 对应最佳臂识别,(ii) 对应遗憾最小化——但同时实现这两个目标仍是一个开放性问题,尽管其具有重要的实际意义。本文提出**遗憾最优最佳臂识别**框架,旨在达成这双重目标。针对预定义停止时间和自适应停止时间两种需求,我们分别提出了名为 EOCP 的算法及其变体。该算法不仅在高斯老虎机和一般老虎机中均能达到渐近最优遗憾,还能在预定义停止时间下以 $\mathcal{O}(\log T)$ 轮、在自适应停止时间下以 $\mathcal{O}(\log^2 T)$ 轮锁定最优臂。我们进一步刻画了 ROBAI 锁定时间(等价于样本复杂度)的下界,表明 EOCP 及其变体在预定义停止时间下是样本最优的,在自适应停止时间下是近乎样本最优的。数值实验结果验证了我们的理论分析,并揭示了一个由经典 UCB 算法带来的有趣“过度探索”现象:尽管 EOCP 比 UCB 更早停止探索(即 $\mathcal{O}(\log T)$ 对比 $\mathcal{O}(T)$),其遗憾却更小,这表明过度探索是不必要的,并可能损害系统性能。