We focus on developing efficient and reliable policy optimization strategies for robot learning with real-world data. In recent years, policy gradient methods have emerged as a promising paradigm for training control policies in simulation. However, these approaches often remain too data inefficient or unreliable to train on real robotic hardware. In this paper we introduce a novel policy gradient-based policy optimization framework which systematically leverages a (possibly highly simplified) first-principles model and enables learning precise control policies with limited amounts of real-world data. Our approach $1)$ uses the derivatives of the model to produce sample-efficient estimates of the policy gradient and $2)$ uses the model to design a low-level tracking controller, which is embedded in the policy class. Theoretical analysis provides insight into how the presence of this feedback controller addresses overcomes key limitations of stand-alone policy gradient methods, while hardware experiments with a small car and quadruped demonstrate that our approach can learn precise control strategies reliably and with only minutes of real-world data.
翻译:我们聚焦于利用真实数据开发高效且可靠的机器人学习策略优化方法。近年来,策略梯度方法已成为在仿真环境中训练控制策略的有前景范式。然而,这些方法在真实机器人硬件上训练时往往数据效率过低或不可靠。本文提出一种新颖的基于策略梯度的策略优化框架,该框架系统性地利用(可能高度简化的)第一性原理模型,能够以有限量的真实世界数据学习精确控制策略。我们的方法:1)利用模型导数生成样本高效的策略梯度估计;2)利用模型设计嵌入到策略类中的低层跟踪控制器。理论分析揭示了该反馈控制器的存在如何克服纯策略梯度方法的关键局限性,而通过小型车辆和四足机器人的硬件实验表明,该方法仅需几分钟的真实数据即可可靠地学习精确控制策略。