Reusing large datasets is crucial to scale vision-based robotic manipulators to everyday scenarios due to the high cost of collecting robotic datasets. However, robotic platforms possess varying control schemes, camera viewpoints, kinematic configurations, and end-effector morphologies, posing significant challenges when transferring manipulation skills from one platform to another. To tackle this problem, we propose a set of key design decisions to train a single policy for deployment on multiple robotic platforms. Our framework first aligns the observation and action spaces of our policy across embodiments via utilizing wrist cameras and a unified, but modular codebase. To bridge the remaining domain shift, we align our policy's internal representations across embodiments through contrastive learning. We evaluate our method on a dataset collected over 60 hours spanning 6 tasks and 3 robots with varying joint configurations and sizes: the WidowX 250S, the Franka Emika Panda, and the Sawyer. Our results demonstrate significant improvements in success rate and sample efficiency for our policy when using new task data collected on a different robot, validating our proposed design decisions. More details and videos can be found on our anonymized project website: https://sites.google.com/view/polybot-multirobot
翻译:摘要:由于收集机器人数据集的成本高昂,重用大规模数据集对于将基于视觉的机械臂操作扩展到日常场景至关重要。然而,机器人平台在控制方案、视角、运动学构型以及末端执行器形态方面存在差异,这给跨平台转移操作技能带来了巨大挑战。为解决这一问题,我们提出了一系列关键设计决策,用于训练一个可在多个机器人平台上部署的通用策略。我们的框架首先通过使用腕部相机和统一但模块化的代码库,对齐策略在不同实体间的观测与动作空间。为弥合剩余域偏移,我们通过对比学习对齐策略在不同实体间的内部表征。我们在一个涵盖6个任务、3种具有不同关节构型与尺寸的机器人(WidowX 250S、Franka Emika Panda 和 Sawyer)的60小时数据集上评估了该方法。结果表明,当使用在不同机器人上收集的新任务数据时,我们的策略在成功率和样本效率上均有显著提升,验证了所提出的设计决策。更多细节与视频请见匿名项目网站:https://sites.google.com/view/polybot-multirobot