This paper is motivated by recent developments in the linear bandit literature, which have revealed a discrepancy between the promising empirical performance of algorithms such as Thompson sampling and Greedy, when compared to their pessimistic theoretical regret bounds. The challenge arises from the fact that while these algorithms may perform poorly in certain problem instances, they generally excel in typical instances. To address this, we propose a new data-driven technique that tracks the geometry of the uncertainty ellipsoid, enabling us to establish an instance-dependent frequentist regret bound for a broad class of algorithms, including Greedy, OFUL, and Thompson sampling. This result empowers us to identify and ``course-correct" instances in which the base algorithms perform poorly. The course-corrected algorithms achieve the minimax optimal regret of order $\tilde{\mathcal{O}}(d\sqrt{T})$, while retaining most of the desirable properties of the base algorithms. We present simulation results to validate our findings and compare the performance of our algorithms with the baselines.
翻译:本文受线性赌博机领域最新研究进展的驱动,该领域揭示了汤普森采样与贪婪算法在实际应用中表现优异,但其理论遗憾界却呈悲观态度的矛盾现象。这一挑战源于此类算法在特定问题实例中表现欠佳,但在典型实例中通常表现出色。为此,我们提出一种追踪不确定性椭球几何形态的新型数据驱动技术,从而为包括贪婪算法、OFUL与汤普森采样在内的广泛算法类别建立实例相关的频率学派遗憾界。该成果使我们能够识别并"纠偏"基础算法表现欠佳的实例。经纠偏后的算法在保持基础算法大部分理想特性的同时,实现了阶数为$\tilde{\mathcal{O}}(d\sqrt{T})$的极小极大最优遗憾。我们通过仿真结果验证了该发现,并将所提出算法与基线模型的性能进行了比较。