Parkour is a grand challenge for legged locomotion that requires robots to overcome various obstacles rapidly in complex environments. Existing methods can generate either diverse but blind locomotion skills or vision-based but specialized skills by using reference animal data or complex rewards. However, autonomous parkour requires robots to learn generalizable skills that are both vision-based and diverse to perceive and react to various scenarios. In this work, we propose a system for learning a single end-to-end vision-based parkour policy of diverse parkour skills using a simple reward without any reference motion data. We develop a reinforcement learning method inspired by direct collocation to generate parkour skills, including climbing over high obstacles, leaping over large gaps, crawling beneath low barriers, squeezing through thin slits, and running. We distill these skills into a single vision-based parkour policy and transfer it to a quadrupedal robot using its egocentric depth camera. We demonstrate that our system can empower two different low-cost robots to autonomously select and execute appropriate parkour skills to traverse challenging real-world environments.
翻译:跑酷是足式运动的一项重大挑战,要求机器人在复杂环境中快速克服各种障碍。现有方法可通过参考动物数据或复杂奖励生成多样化但缺乏视觉感知的运动技能,或基于视觉但高度特化的技能。然而,自主跑酷要求机器人学习兼具视觉感知与多样性的可泛化技能,以感知并应对各种场景。本文提出一种无需任何参考运动数据,仅通过简单奖励即可学习单一端到端视觉跑酷策略的系统,该策略可执行多种跑酷技能。我们开发了一种受直接配点法启发的强化学习方法,用于生成包括翻越障碍物、跨越宽缝隙、钻过低矮屏障、挤过狭窄缝隙及奔跑等跑酷技能。我们将这些技能蒸馏至单一视觉跑酷策略中,并通过机器人的自视深度相机将其迁移至四足机器人平台。实验表明,本系统可赋能两种低成本机器人自主选择并执行相应跑酷技能,在真实复杂环境中完成穿越任务。