Model selection in the context of bandit optimization is a challenging problem, as it requires balancing exploration and exploitation not only for action selection, but also for model selection. One natural approach is to rely on online learning algorithms that treat different models as experts. Existing methods, however, scale poorly ($\text{poly}M$) with the number of models $M$ in terms of their regret. Our key insight is that, for model selection in linear bandits, we can emulate full-information feedback to the online learner with a favorable bias-variance trade-off. This allows us to develop ALEXP, which has an exponentially improved ($\log M$) dependence on $M$ for its regret. ALEXP has anytime guarantees on its regret, and neither requires knowledge of the horizon $n$, nor relies on an initial purely exploratory stage. Our approach utilizes a novel time-uniform analysis of the Lasso, establishing a new connection between online learning and high-dimensional statistics.
翻译:在强盗优化中进行模型选择是一项具有挑战性的问题,因为它不仅需要在动作选择上平衡探索与利用,还需要在模型选择上进行权衡。一种自然的方法是依赖将不同模型视为专家的在线学习算法。然而,现有方法在遗憾值方面对模型数量 $M$ 的扩展性较差($\text{poly}M$)。我们的关键见解是,对于线性强盗中的模型选择,我们可以通过有利的偏差-方差权衡,向在线学习器模拟完全信息反馈。这使我们能够开发出 ALEXP,其在遗憾值上对 $M$ 的依赖呈指数级改进($\log M$)。ALEXP 具有关于遗憾值的任意时刻保证,既不需要知道时间范围 $n$,也不依赖于初始的纯探索阶段。我们的方法利用了对 Lasso 的新颖时间一致性分析,建立了在线学习与高维统计之间的新联系。