While most recent advancements in legged robot control have been driven by model-free reinforcement learning, we explore the potential of differentiable simulation. Differentiable simulation promises faster convergence and more stable training by computing low-variant first-order gradients using the robot model, but so far, its use for legged robot control has remained limited to simulation. The main challenge with differentiable simulation lies in the complex optimization landscape of robotic tasks due to discontinuities in contact-rich environments, e.g., quadruped locomotion. This work proposes a new, differentiable simulation framework to overcome these challenges. The key idea involves decoupling the complex whole-body simulation, which may exhibit discontinuities due to contact, into two separate continuous domains. Subsequently, we align the robot state resulting from the simplified model with a more precise, non-differentiable simulator to maintain sufficient simulation accuracy. Our framework enables learning quadruped walking in minutes using a single simulated robot without any parallelization. When augmented with GPU parallelization, our approach allows the quadruped robot to master diverse locomotion skills, including trot, pace, bound, and gallop, on challenging terrains in minutes. Additionally, our policy achieves robust locomotion performance in the real world zero-shot. To the best of our knowledge, this work represents the first demonstration of using differentiable simulation for controlling a real quadruped robot. This work provides several important insights into using differentiable simulations for legged locomotion in the real world.
翻译:尽管近期腿足机器人控制的大多数进展由无模型强化学习驱动,但我们探索了可微分仿真的潜力。可微分仿真通过利用机器人模型计算低方差一阶梯度,有望实现更快的收敛速度和更稳定的训练,但迄今为止,其在腿足机器人控制中的应用仍局限于仿真环境。可微分仿真的主要挑战在于:接触丰富的环境(如四足运动)中的不连续性会导致机器人任务的优化景观复杂化。本文提出一种新的可微分仿真框架以克服这些挑战。其核心思想是将可能因接触产生不连续性的复杂全身仿真解耦为两个独立的连续域,随后将简化模型生成的机器人状态与更精确的不可微分仿真器对齐,以维持足够的仿真精度。该框架使单个仿真机器人无需任何并行化即可在数分钟内学会四步行走。当结合GPU并行化后,我们的方法能使四足机器人在数分钟内掌握多种步态技能(包括小跑、踱步、跳跃和飞奔),并适应复杂地形。此外,我们的策略在真实世界中实现了零样本鲁棒运动性能。据我们所知,本研究首次证明了可微分仿真在控制真实四足机器人中的应用,并为在现实世界中利用可微分仿真实现腿足运动提供了若干重要见解。