Imitation can allow us to quickly gain an understanding of a new task. Through a demonstration, we can gain direct knowledge about which actions need to be performed and which goals they have. In this paper, we introduce a new approach to imitation learning that tackles the challenges of a robot imitating a human, such as the change in perspective and body schema. Our approach can use a single human demonstration to abstract information about the demonstrated task, and use that information to generalise and replicate it. We facilitate this ability by a new integration of two state-of-the-art methods: a diffusion action segmentation model to abstract temporal information from the demonstration and an open vocabulary object detector for spatial information. Furthermore, we refine the abstracted information and use symbolic reasoning to create an action plan utilising inverse kinematics, to allow the robot to imitate the demonstrated action.
翻译:模仿能够帮助我们快速理解一项新任务。通过示范,我们可以直接获取关于需要执行哪些动作及其目标的知识。本文提出了一种新的模仿学习方法,以应对机器人模仿人类时所面临的挑战,如视角和身体图式的变化。我们的方法可利用单次人类示范来提取关于所演示任务的抽象信息,并将这些信息用于归纳和复现任务。我们通过整合两种当前最先进的技术来增强这一能力:一种扩散动作分割模型,用于从示范中提取时间信息;以及一种开放词汇目标检测器,用于提取空间信息。此外,我们对提取的信息进行精炼,并利用符号推理结合逆运动学创建动作规划,从而使机器人能够模仿演示的动作。