This letter presents a versatile control method for dynamic and robust legged locomotion that integrates model-based optimal control with reinforcement learning (RL). Our approach involves training an RL policy to imitate reference motions generated on-demand through solving a finite-horizon optimal control problem. This integration enables the policy to leverage human expertise in generating motions to imitate while also allowing it to generalize to more complex scenarios that require a more complex dynamics model. Our method successfully learns control policies capable of generating diverse quadrupedal gait patterns and maintaining stability against unexpected external perturbations in both simulation and hardware experiments. Furthermore, we demonstrate the adaptability of our method to more complex locomotion tasks on uneven terrain without the need for excessive reward shaping or hyperparameter tuning.
翻译:本论文提出一种将基于模型的最优控制与强化学习(RL)相结合的多功能控制方法,用于实现动态且鲁棒的腿式运动。我们的方法通过训练一个RL策略,模仿通过求解有限时域最优控制问题按需生成的参考运动。这种整合使策略既能利用人类在生成模仿运动方面的专业知识,又能使其泛化到需要更复杂动力学模型的复杂场景中。我们的方法成功学习到能够生成多种四足步态模式的控制策略,并在仿真和硬件实验中保持对外部意外扰动的稳定性。此外,我们证明了该方法在非平坦地形上完成更复杂运动任务时的适应性,而无需进行过多的奖励塑造或超参数调整。