To succeed in the real world, robots must cope with situations that differ from those seen during training. We study the problem of adapting on-the-fly to such novel scenarios during deployment, by drawing upon a diverse repertoire of previously learned behaviors. Our approach, RObust Autonomous Modulation (ROAM), introduces a mechanism based on the perceived value of pre-trained behaviors to select and adapt pre-trained behaviors to the situation at hand. Crucially, this adaptation process all happens within a single episode at test time, without any human supervision. We provide theoretical analysis of our selection mechanism and demonstrate that ROAM enables a robot to adapt rapidly to changes in dynamics both in simulation and on a real Go1 quadruped, even successfully moving forward with roller skates on its feet. Our approach adapts over 2x as efficiently compared to existing methods when facing a variety of out-of-distribution situations during deployment by effectively choosing and adapting relevant behaviors on-the-fly.
翻译:要在现实世界中成功部署,机器人必须应对与训练时不同的场景。我们研究如何利用先前学习的多样化行为库,在部署过程中即时适应这些新场景。我们提出的鲁棒自主行为调制方法(ROAM)引入了一种基于预训练行为感知价值的机制,能够根据当前情境选择并调整预训练行为。关键在于,整个适应过程在测试阶段的单一回合内完成,无需任何人工干预。我们对选择机制进行了理论分析,并证明ROAM可使机器人在仿真环境和真实Go1四足机器人上快速适应动力学变化——即使脚部安装轮滑鞋也能成功前行。面对部署期间出现的各类分布外场景时,我们的方法通过有效实时选择和调整相关行为,适应效率比现有方法提升2倍以上。