We derive the first finite-time logarithmic regret bounds for Bayesian bandits. For Gaussian bandits, we obtain a $O(c_h \log^2 n)$ bound, where $c_h$ is a prior-dependent constant. This matches the asymptotic lower bound of Lai (1987). Our proofs mark a technical departure from prior works, and are simple and general. To show generality, we apply our technique to linear bandits. Our bounds shed light on the value of the prior in the Bayesian setting, both in the objective and as a side information given to the learner. They significantly improve the $\tilde{O}(\sqrt{n})$ bounds, that despite the existing lower bounds, have become standard in the literature.
翻译:我们首次推导了贝叶斯赌博机问题的有限时间对数后悔界。对于高斯赌博机,我们得到了一个$O(c_h \log^2 n)$的界,其中$c_h$是一个依赖于先验的常数。这与Lai(1987)的渐近下界相匹配。我们的证明在技术上不同于先前的工作,且简洁而通用。为展示其通用性,我们将该方法应用于线性赌博机。该界揭示了贝叶斯框架中先验的价值——无论从目标函数角度,还是作为提供给学习者的辅助信息。它显著改进了$\tilde{O}(\sqrt{n})$界,尽管存在已知的下界,后者已成为文献中的标准结果。