We investigate the extent to which offline demonstration data can improve online learning. It is natural to expect some improvement, but the question is how, and by how much? We show that the degree of improvement must depend on the quality of the demonstration data. To generate portable insights, we focus on Thompson sampling (TS) applied to a multi-armed bandit as a prototypical online learning algorithm and model. The demonstration data is generated by an expert with a given competence level, a notion we introduce. We propose an informed TS algorithm that utilizes the demonstration data in a coherent way through Bayes' rule and derive a prior-dependent Bayesian regret bound. This offers insight into how pretraining can greatly improve online performance and how the degree of improvement increases with the expert's competence level. We also develop a practical, approximate informed TS algorithm through Bayesian bootstrapping and show substantial empirical regret reduction through experiments.
翻译:我们研究了离线演示数据能在多大程度上改善在线学习。虽然预期会有一定程度的提升,但问题在于如何提升以及提升多少。研究表明,提升程度必须取决于演示数据的质量。为得出具有普适性的见解,我们以多臂老虎机上的汤普森采样(TS)作为典型的在线学习算法和模型进行研究。演示数据由具有特定胜任力水平(本文引入的概念)的专家生成。我们提出了一种有依据的TS算法,通过贝叶斯规则以连贯的方式利用演示数据,并推导出与先验相关的贝叶斯遗憾上界。这一结果揭示了预训练如何显著提升在线性能,以及提升幅度如何随专家胜任力水平的增加而增大。我们还通过贝叶斯自助法开发了一种实用且近似的有依据TS算法,并通过实验证明了其能够大幅降低经验遗憾。