Recent advances in deep reinforcement learning (RL) based techniques combined with training in simulation have offered a new approach to developing robust controllers for legged robots. However, the application of such approaches to real hardware has largely been limited to quadrupedal robots with direct-drive actuators and light-weight bipedal robots with low gear-ratio transmission systems. Application to real, life-sized humanoid robots has been less common arguably due to a large sim2real gap. In this paper, we present an approach for effectively overcoming the sim2real gap issue for humanoid robots arising from inaccurate torque-tracking at the actuator level. Our key idea is to utilize the current feedback from the actuators on the real robot, after training the policy in a simulation environment artificially degraded with poor torque-tracking. Our approach successfully trains a unified, end-to-end policy in simulation that can be deployed on a real HRP-5P humanoid robot to achieve bipedal locomotion. Through ablations, we also show that a feedforward policy architecture combined with targeted dynamics randomization is sufficient for zero-shot sim2real success, thus eliminating the need for computationally expensive, memory-based network architectures. Finally, we validate the robustness of the proposed RL policy by comparing its performance against a conventional model-based controller for walking on uneven terrain with the real robot.
翻译:基于深度强化学习(RL)技术结合仿真训练的最新进展,为开发腿式机器人的鲁棒控制器提供了新思路。然而,此类方法在实际硬件中的应用主要局限于配备直驱执行器的四足机器人,以及采用低减速比传动系统的轻型双足机器人。由于模拟到现实(sim2real)差距的显著存在,该类方法在真实尺寸人形机器人上的应用尚不常见。本文提出了一种有效克服人形机器人因执行器层面扭矩跟踪不准确而导致的sim2real差距问题的方案。核心思想是:在人为引入扭矩跟踪误差劣化效果的仿真环境中完成策略训练后,利用真实机器人执行器的电流反馈。该方法成功在仿真环境中训练出统一的端到端策略,并使其在真实HRP-5P人形机器人上实现双足行走。通过消融实验证明,前馈策略架构结合目标性动力学随机化足以实现零样本sim2real迁移,从而避免采用计算成本高昂的基于记忆的网络架构。最后,通过真实机器人在不平坦地形上的行走实验,将所提强化学习策略与传统基于模型的控制器进行性能对比,验证了其鲁棒性。