Follow-the-Regularized-Leader (FTRL) is a powerful framework for various online learning problems. By designing its regularizer and learning rate to be adaptive to past observations, FTRL is known to work adaptively to various properties of an underlying environment. However, most existing adaptive learning rates are for online learning problems with a minimax regret of $\Theta(\sqrt{T})$ for the number of rounds $T$, and there are only a few studies on adaptive learning rates for problems with a minimax regret of $\Theta(T^{2/3})$, which include several important problems dealing with indirect feedback. To address this limitation, we establish a new adaptive learning rate framework for problems with a minimax regret of $\Theta(T^{2/3})$. Our learning rate is designed by matching the stability, penalty, and bias terms that naturally appear in regret upper bounds for problems with a minimax regret of $\Theta(T^{2/3})$. As applications of this framework, we consider two major problems dealing with indirect feedback: partial monitoring and graph bandits. We show that FTRL with our learning rate and the Tsallis entropy regularizer improves existing Best-of-Both-Worlds (BOBW) regret upper bounds, which achieve simultaneous optimality in the stochastic and adversarial regimes. The resulting learning rate is surprisingly simple compared to the existing learning rates for BOBW algorithms for problems with a minimax regret of $\Theta(T^{2/3})$.
翻译:跟随正则化领导者(FTRL)是解决各类在线学习问题的强大框架。通过设计其正则化项和学习率,使其能够自适应于过去的观测结果,FTRL 已知能够自适应地应对底层环境的各种特性。然而,现有的大多数自适应学习率都是针对极小极大遗憾为 $\Theta(\sqrt{T})$(其中 $T$ 为回合数)的在线学习问题,而对于极小极大遗憾为 $\Theta(T^{2/3})$ 的问题(其中包括几个处理间接反馈的重要问题)的自适应学习率研究则很少。为了弥补这一不足,我们为极小极大遗憾为 $\Theta(T^{2/3})$ 的问题建立了一个新的自适应学习率框架。我们的学习率是通过匹配在极小极大遗憾为 $\Theta(T^{2/3})$ 的问题的遗憾上界中自然出现的稳定性项、惩罚项和偏差项来设计的。作为该框架的应用,我们考虑了两个处理间接反馈的主要问题:部分监测和图赌博机。我们证明了,采用我们的学习率和 Tsallis 熵正则化器的 FTRL 改进了现有的“两全其美”(BOBW)遗憾上界,该上界在随机和对抗两种机制下同时实现了最优性。与现有的针对极小极大遗憾为 $\Theta(T^{2/3})$ 问题的 BOBW 算法的学习率相比,所得的学习率出人意料地简单。