Learning from demonstrations (LfD) methods guide learning agents to a desired solution using demonstrations from a teacher. While some LfD methods can handle small mismatches in the action spaces of the teacher and student, here we address the case where the teacher demonstrates the task in an action space that can be substantially different from that of the student -- thereby inducing a large action space mismatch. We bridge this gap with a framework, Morphological Adaptation in Imitation Learning (MAIL), that allows training an agent from demonstrations by other agents with significantly different morphologies (from the student or each other). MAIL is able to learn from suboptimal demonstrations, so long as they provide some guidance towards a desired solution. We demonstrate MAIL on challenging household cloth manipulation tasks and introduce a new DRY CLOTH task -- cloth manipulation in 3D task with obstacles. In these tasks, we train a visual control policy for a robot with one end-effector using demonstrations from a simulated agent with two end-effectors. MAIL shows up to 27% improvement over LfD and non-LfD baselines. It is deployed to a real Franka Panda robot, and can handle multiple variations in cloth properties (color, thickness, size, material) and pose (rotation and translation). We further show generalizability to transfers from n-to-m end-effectors, in the context of a simple rearrangement task.
翻译:基于示教学习(LfD)方法通过教师提供的示范引导学习智能体达到期望解。虽然部分LfD方法能处理教师与学生动作空间存在较小差异的情况,但本文聚焦于教师在与学生显著不同的动作空间中演示任务的场景——由此造成巨大的动作空间失配。我们提出"模仿学习中的形态适应"(MAIL)框架来弥合这一鸿沟,该框架允许智能体通过其他具有不同形态(与学生或彼此之间)的智能体的示范进行训练。MAIL能够从次优示范中学习,只要这些示范能为目标解提供一定引导。我们在具有挑战性的家庭布料操控任务中验证MAIL,并引入新的DRY CLOTH任务——含障碍物的三维布料操控任务。在这些任务中,我们利用双末端执行器模拟智能体的示范,为单末端执行器机器人训练视觉控制策略。MAIL相比LfD与非LfD基线方法实现了最高27%的性能提升。该框架已部署至真实Franka Panda机器人,能处理布料属性(颜色、厚度、尺寸、材质)和位姿(旋转与平移)的多样变化。我们进一步在简单重排任务背景下,展示了从n到m个末端执行器迁移的泛化能力。