The linear bandit problem has been studied for many years in both stochastic and adversarial settings. Designing an algorithm that can optimize the environment without knowing the loss type attracts lots of interest. \citet{LeeLWZ021} propose an algorithm that actively detects the loss type and then switches between different algorithms specially designed for different settings. However, such an approach requires meticulous designs to perform well in all settings. Follow-the-regularized-leader (FTRL) is another popular algorithm type that can adapt to different environments. This algorithm is of simple design and the regret bounds are shown to be optimal in traditional multi-armed bandit problems compared with the detect-switch type algorithms. Designing an FTRL-type algorithm for linear bandits is an important question that has been open for a long time. In this paper, we prove that the FTRL-type algorithm with a negative entropy regularizer can achieve the best-of-three-world results for the linear bandit problem with the tacit cooperation between the choice of the learning rate and the specially designed self-bounding inequality.
翻译:线性赌博机问题在随机和对抗两种环境中已被研究多年。设计一种无需知晓损失类型即可优化环境的算法吸引了广泛关注。Lee等人(2020)提出了一种主动检测损失类型并在不同场景下切换专门设计算法的方案。然而,这种方法需要精细的设计才能在所有环境中表现良好。跟随正则化领导者(FTRL)是另一种能适应不同环境的流行算法类型。该算法设计简洁,且与传统检测切换型算法相比,其在经典多臂赌博机问题中的遗憾界已被证明达到最优。设计适用于线性赌博机的FTRL型算法是一个长期悬而未决的重要问题。本文证明,通过学习率的选择与专门构建的自限不等式之间的默契配合,基于负熵正则化的FTRL型算法能够在三种环境中均获得最优结果。