We study online learning with an additional offline dataset in the stochastic linear bandit setting. Although this problem arises frequently in practice, the offline-to-online tradeoff remains poorly understood in structured environments. We propose a linear bandit algorithm that balances this tradeoff: it relies on offline data during early rounds, and increasingly favors exploration as the horizon grows. We establish regret bounds showing that our method is simultaneously competitive with both purely online and purely offline solutions. In particular, it achieves sublinear regret relative to the optimal action in the number of online interactions, while its regret relative to an offline reference decreases as the number of offline samples grows. Empirical results further demonstrate its effectiveness across various problem parameters.
翻译:我们在随机线性赌博机设置中研究带有额外离线数据集的在线学习问题。尽管该问题在实践中频繁出现,但在结构化环境中,离线到在线的权衡仍未被充分理解。我们提出了一种线性赌博机算法来平衡这种权衡:该算法在早期轮次依赖于离线数据,并随着时间范围的扩大逐渐偏向于探索。我们建立了遗憾界,表明我们的方法同时与纯在线和纯离线解决方案具有竞争力。特别地,该方法相对于最优动作的在线交互次数实现了次线性遗憾,同时相对于离线参考的遗憾随着离线样本数量的增加而减少。实验结果进一步证明了该方法在各种问题参数下的有效性。