Model-based approaches to reinforcement learning (MBRL) exhibit favorable performance in practice, but their theoretical guarantees in large spaces are mostly restricted to the setting when transition model is Gaussian or Lipschitz, and demands a posterior estimate whose representational complexity grows unbounded with time. In this work, we develop a novel MBRL method (i) which relaxes the assumptions on the target transition model to belong to a generic family of mixture models; (ii) is applicable to large-scale training by incorporating a compression step such that the posterior estimate consists of a Bayesian coreset of only statistically significant past state-action pairs; and (iii) exhibits a sublinear Bayesian regret. To achieve these results, we adopt an approach based upon Stein's method, which, under a smoothness condition on the constructed posterior and target, allows distributional distance to be evaluated in closed form as the kernelized Stein discrepancy (KSD). The aforementioned compression step is then computed in terms of greedily retaining only those samples which are more than a certain KSD away from the previous model estimate. Experimentally, we observe that this approach is competitive with several state-of-the-art RL methodologies, and can achieve up-to 50 percent reduction in wall clock time in some continuous control environments.
翻译:基于模型的强化学习方法在实践中表现出优越的性能,但其在大空间下的理论保证主要局限于转移模型为高斯或李普希茨的情形,且要求后验估计的表示复杂度随时间无限增长。本文开发了一种新颖的基于模型的强化学习方法:(i) 放宽了目标转移模型属于一般混合模型族的假设;(ii) 通过引入压缩步骤适用于大规模训练,使得后验估计仅包含统计上显著的过去状态-动作对的贝叶斯核心集;(iii) 具有次线性贝叶斯遗憾。为实现这些结果,我们采用了基于斯坦因方法的技术路径,在构建的后验与目标满足光滑性条件下,可闭式计算分布距离为核化斯坦因差异。上述压缩步骤通过贪婪保留那些与先前模型估计的核化斯坦因差异超过特定阈值的样本实现。实验表明,该方法与多种最先进的强化学习技术具有竞争力,在部分连续控制环境中可实现高达50%的墙钟时间缩减。