Sim-to-real is a mainstream method to cope with the large number of trials needed by typical deep reinforcement learning methods. However, transferring a policy trained in simulation to actual hardware remains an open challenge due to the reality gap. In particular, the characteristics of actuators in legged robots have a considerable influence on sim-to-real transfer. There are two challenges: 1) High reduction ratio gears are widely used in actuators, and the reality gap issue becomes especially pronounced when backdrivability is considered in controlling joints compliantly. 2) The difficulty in achieving stable bipedal locomotion causes typical system identification methods to fail to sufficiently transfer the policy. For these two challenges, we propose 1) a new simulation model of gears and 2) a method for system identification that can utilize failed attempts. The method's effectiveness is verified using a biped robot, the ROBOTIS-OP3, and the sim-to-real transferred policy can stabilize the robot under severe disturbances and walk on uneven surfaces without using force and torque sensors.
翻译:仿真实体迁移是应对典型深度强化学习方法所需大量试错的主流手段。然而,由于现实差异的存在,将仿真训练的策略迁移至实际硬件仍面临开放挑战。其中,足式机器人执行器特性对仿真实体迁移具有显著影响。现有两大挑战:1)执行器广泛采用高减速比齿轮,而柔顺关节控制中反向驱动特性加剧了现实差异问题;2)双足稳定运动的实现难度导致常规系统辨识方法无法充分迁移策略。针对上述挑战,本文提出:1)新型齿轮仿真模型;2)可有效利用失败尝试的系统辨识方法。通过双足机器人ROBOTIS-OP3验证该方法有效性,证实经仿真实体迁移后的策略能在无力和扭矩传感器条件下,有效维持机器人在剧烈扰动下的稳定姿态,并实现非平坦地形行走。