Training long-horizon robotic policies in complex physical environments is essential for many applications, such as robotic manipulation. However, learning a policy that can generalize to unseen tasks is challenging. In this work, we propose to achieve one-shot task generalization by decoupling plan generation and plan execution. Specifically, our method solves complex long-horizon tasks in three steps: build a paired abstract environment by simplifying geometry and physics, generate abstract trajectories, and solve the original task by an abstract-to-executable trajectory translator. In the abstract environment, complex dynamics such as physical manipulation are removed, making abstract trajectories easier to generate. However, this introduces a large domain gap between abstract trajectories and the actual executed trajectories as abstract trajectories lack low-level details and are not aligned frame-to-frame with the executed trajectory. In a manner reminiscent of language translation, our approach leverages a seq-to-seq model to overcome the large domain gap between the abstract and executable trajectories, enabling the low-level policy to follow the abstract trajectory. Experimental results on various unseen long-horizon tasks with different robot embodiments demonstrate the practicability of our methods to achieve one-shot task generalization.
翻译:训练在复杂物理环境中的长期机器人策略对于许多应用(如机器人操作)至关重要。然而,学习能够泛化到未见任务的策略具有挑战性。在本工作中,我们提出通过解耦计划生成与计划执行来实现一次性任务泛化。具体而言,我们的方法通过三个步骤解决复杂的长期任务:通过简化几何与物理构建配对的抽象环境、生成抽象轨迹,并通过抽象到可执行轨迹翻译器解决原始任务。在抽象环境中,复杂动力学(如物理操作)被移除,使得抽象轨迹更容易生成。然而,这导致了抽象轨迹与实际执行轨迹之间存在巨大的领域差距,因为抽象轨迹缺乏低级细节,且未与执行轨迹进行逐帧对齐。类似于语言翻译,我们的方法利用序列到序列模型克服抽象轨迹与可执行轨迹之间的巨大领域差距,使低级策略能够遵循抽象轨迹。在不同机器人形态的多种未见长期任务上的实验结果证明了我们的方法在实现一次性任务泛化方面的实用性。