In this paper, we study the problem of efficient online reinforcement learning in the infinite horizon setting when there is an offline dataset to start with. We assume that the offline dataset is generated by an expert but with unknown level of competence, i.e., it is not perfect and not necessarily using the optimal policy. We show that if the learning agent models the behavioral policy (parameterized by a competence parameter) used by the expert, it can do substantially better in terms of minimizing cumulative regret, than if it doesn't do that. We establish an upper bound on regret of the exact informed PSRL algorithm that scales as $\tilde{O}(\sqrt{T})$. This requires a novel prior-dependent regret analysis of Bayesian online learning algorithms for the infinite horizon setting. We then propose an approximate Informed RLSVI algorithm that we can interpret as performing imitation learning with the offline dataset, and then performing online learning.
翻译:本文研究了在存在离线数据集启动的情况下,无限时域环境中高效在线强化学习的问题。我们假设离线数据由专家生成,但该专家的能力水平未知,即数据并非完美,且未必采用最优策略。我们证明,若学习智能体对专家使用的行为策略(通过能力参数进行建模)进行建模,其在最小化累积遗憾方面的表现将显著优于不进行此类建模的情况。我们为精确的知情PSRL算法建立了遗憾的上界,该上界量级为$\tilde{O}(\sqrt{T})$。这需要对无限时域环境下的贝叶斯在线学习算法进行一种新颖的先验依赖遗憾分析。随后,我们提出了一种近似的知情RLSVI算法,该算法可被解释为首先利用离线数据集进行模仿学习,随后进行在线学习。