Designing and deriving effective model-based reinforcement learning (MBRL) algorithms with a performance improvement guarantee is challenging, mainly attributed to the high coupling between model learning and policy optimization. Many prior methods that rely on return discrepancy to guide model learning ignore the impacts of model shift, which can lead to performance deterioration due to excessive model updates. Other methods use performance difference bound to explicitly consider model shift. However, these methods rely on a fixed threshold to constrain model shift, resulting in a heavy dependence on the threshold and a lack of adaptability during the training process. In this paper, we theoretically derive an optimization objective that can unify model shift and model bias and then formulate a fine-tuning process. This process adaptively adjusts the model updates to get a performance improvement guarantee while avoiding model overfitting. Based on these, we develop a straightforward algorithm USB-PO (Unified model Shift and model Bias Policy Optimization). Empirical results show that USB-PO achieves state-of-the-art performance on several challenging benchmark tasks.
翻译:设计和推导具有性能改进保证的基于模型的强化学习算法颇具挑战性,主要归因于模型学习与策略优化之间的高度耦合。许多先前依赖回报差异指导模型学习的方法忽视了模型偏移的影响,这可能导致因过度模型更新而发生性能恶化;其他方法则利用性能差异上界显式考虑模型偏移,但这些方法依赖固定阈值约束模型偏移,导致对阈值的重度依赖且在训练过程中缺乏自适应性。本文从理论上推导了一个可统一模型偏移与模型偏差的优化目标,并由此构建了一个微调过程。该过程能自适应调整模型更新,在保证性能提升的同时避免模型过拟合。基于此,我们提出了一种简洁算法USB-PO(统一模型偏移与模型偏差的策略优化)。实验结果表明,USB-PO在多个具有挑战性的基准测试任务上达到了当前最优性能。