Offline reinforcement learning (offline RL) considers problems where learning is performed using only previously collected samples and is helpful for the settings in which collecting new data is costly or risky. In model-based offline RL, the learner performs estimation (or optimization) using a model constructed according to the empirical transition frequencies. We analyze the sample complexity of vanilla model-based offline RL with dependent samples in the infinite-horizon discounted-reward setting. In our setting, the samples obey the dynamics of the Markov decision process and, consequently, may have interdependencies. Under no assumption of independent samples, we provide a high-probability, polynomial sample complexity bound for vanilla model-based off-policy evaluation that requires partial or uniform coverage. We extend this result to the off-policy optimization under uniform coverage. As a comparison to the model-based approach, we analyze the sample complexity of off-policy evaluation with vanilla importance sampling in the infinite-horizon setting. Finally, we provide an estimator that outperforms the sample-mean estimator for almost deterministic dynamics that are prevalent in reinforcement learning.
翻译:离线强化学习(离线RL)考虑利用仅先前收集的样本进行学习的问题,在收集新数据成本高昂或具有风险的环境中十分有用。在模型驱动的离线RL中,学习器根据经验转移频率构建的模型进行估计(或优化)。我们分析了无限时域折扣奖励设置下,依赖样本的朴素模型驱动离线RL的样本复杂度。在我们的设定中,样本遵循马尔可夫决策过程的动态特性,因此可能存在相互依赖。在无独立样本假设的条件下,我们为需要部分或均匀覆盖的朴素模型驱动离线策略评估提供了高概率、多项式的样本复杂度界。并将这一结果推广至均匀覆盖下的离线策略优化。作为与模型驱动方法的对比,我们分析了无限时域设置下朴素重要性采样离线策略评估的样本复杂度。最后,我们提出一种在强化学习中常见的几乎确定性动态环境下优于样本均值估计量的估计器。