We propose a novel model-based reinforcement learning algorithm -- Dynamics Learning and predictive control with Parameterized Actions (DLPA) -- for Parameterized Action Markov Decision Processes (PAMDPs). The agent learns a parameterized-action-conditioned dynamics model and plans with a modified Model Predictive Path Integral control. We theoretically quantify the difference between the generated trajectory and the optimal trajectory during planning in terms of the value they achieved through the lens of Lipschitz Continuity. Our empirical results on several standard benchmarks show that our algorithm achieves superior sample efficiency and asymptotic performance than state-of-the-art PAMDP methods.
翻译:我们提出了一种新颖的基于模型的强化学习算法——参数化动作动态学习与预测控制算法(DLPA),用于处理参数化动作马尔可夫决策过程(PAMDPs)。智能体学习一个以参数化动作为条件的动态模型,并结合改进的模型预测路径积分控制进行规划。我们从理论上量化了规划过程中生成轨迹与最优轨迹之间在Lipschitz连续性框架下所实现的价值差异。在多个标准基准测试中的实验结果表明,该算法相较于当前最先进的PAMDP方法,在样本效率和渐近性能上均具有显著优势。