Smoothed online learning has emerged as a popular framework to mitigate the substantial loss in statistical and computational complexity that arises when one moves from classical to adversarial learning. Unfortunately, for some spaces, it has been shown that efficient algorithms suffer an exponentially worse regret than that which is minimax optimal, even when the learner has access to an optimization oracle over the space. To mitigate that exponential dependence, this work introduces a new notion of complexity, the generalized bracketing numbers, which marries constraints on the adversary to the size of the space, and shows that an instantiation of Follow-the-Perturbed-Leader can attain low regret with the number of calls to the optimization oracle scaling optimally with respect to average regret. We then instantiate our bounds in several problems of interest, including online prediction and planning of piecewise continuous functions, which has many applications in fields as diverse as econometrics and robotics.
翻译:平滑在线学习已成为一种流行框架,旨在缓解从经典学习转向对抗性学习时统计和计算复杂度的大幅增加。不幸的是,对于某些空间,即使学习者能够访问该空间上的优化Oracle,高效算法所遭受的遗憾也指数级地劣于极小极大最优值。为缓解这种指数依赖性,本文提出了一种新的复杂度概念——广义括号数,该概念将对抗者的约束与空间大小相结合,并表明遵循扰动领导者算法能够实现低遗憾,且对优化Oracle的调用次数在平均遗憾方面达到最优尺度。随后,我们将我们的界限实例化到若干重要问题中,包括分段连续函数的在线预测与规划,这些问题在计量经济学和机器人学等众多领域具有广泛应用。