Legged robots have enormous potential in their range of capabilities, from navigating unstructured terrains to high-speed running. However, designing robust controllers for highly agile dynamic motions remains a substantial challenge for roboticists. Reinforcement learning (RL) offers a promising data-driven approach for automatically training such controllers. However, exploration in these high-dimensional, underactuated systems remains a significant hurdle for enabling legged robots to learn performant, naturalistic, and versatile agility skills. We propose a framework for training complex robotic skills by transferring experience from existing controllers to jumpstart learning new tasks. To leverage controllers we can acquire in practice, we design this framework to be flexible in terms of their source -- that is, the controllers may have been optimized for a different objective under different dynamics, or may require different knowledge of the surroundings -- and thus may be highly suboptimal for the target task. We show that our method enables learning complex agile jumping behaviors, navigating to goal locations while walking on hind legs, and adapting to new environments. We also demonstrate that the agile behaviors learned in this way are graceful and safe enough to deploy in the real world.
翻译:腿足机器人在从非结构化地形导航到高速奔跑等能力范围内具有巨大潜力。然而,为高度敏捷的动态运动设计鲁棒控制器仍是机器人领域面临的一项重大挑战。强化学习提供了一种有前景的数据驱动方法,可用于自动训练此类控制器。然而,在这些高维、欠驱动系统中的探索仍是一个关键障碍,阻碍了腿足机器人学习高性能、自然且多功能的敏捷技能。我们提出了一种通过从现有控制器迁移经验来快速启动新任务学习、从而训练复杂机器人技能的框架。为了充分利用实践中可获得的控制器,我们设计该框架对其来源具有灵活性——即这些控制器可能针对不同动力学条件下的不同目标进行了优化,或可能需要对环境的不同认知——因而对于目标任务可能高度次优。我们证明,该方法能够学习复杂的敏捷跳跃行为、以后肢行走方式导航至目标位置,并适应新环境。我们还展示了以这种方式习得的敏捷行为足够优雅且安全,可在现实世界中部署。