Visual imitation learning has achieved impressive progress in learning unimanual manipulation tasks from a small set of visual observations, thanks to the latest advances in computer vision. However, learning bimanual coordination strategies and complex object relations from bimanual visual demonstrations, as well as generalizing them to categorical objects in novel cluttered scenes remain unsolved challenges. In this paper, we extend our previous work on keypoints-based visual imitation learning (\mbox{K-VIL})~\cite{gao_kvil_2023} to bimanual manipulation tasks. The proposed Bi-KVIL jointly extracts so-called \emph{Hybrid Master-Slave Relationships} (HMSR) among objects and hands, bimanual coordination strategies, and sub-symbolic task representations. Our bimanual task representation is object-centric, embodiment-independent, and viewpoint-invariant, thus generalizing well to categorical objects in novel scenes. We evaluate our approach in various real-world applications, showcasing its ability to learn fine-grained bimanual manipulation tasks from a small number of human demonstration videos. Videos and source code are available at https://sites.google.com/view/bi-kvil.
翻译:摘要:得益于计算机视觉的最新进展,视觉模仿学习在从少量视觉观测中学习单手操作任务方面取得了显著进展。然而,从双手视觉演示中学习双手协调策略和复杂物体关系,并将其泛化至新型杂乱场景中的类别物体,仍是未解决的挑战。本文在前期基于关键点的视觉模仿学习(K-VIL)~\cite{gao_kvil_2023} 工作基础上,将其扩展至双手操作任务。所提出的 Bi-KVIL 方法联合提取物体与手之间的所谓“混合主从关系”(HMSR)、双手协调策略以及子符号任务表示。我们的双手任务表示以物体为中心、与具体执行体无关且具有视角不变性,因此能良好地泛化至新型场景中的类别物体。我们在多种真实世界应用中评估了该方法,展示了其从少量人类演示视频中学习精细双手操作任务的能力。视频和源代码请参见 https://sites.google.com/view/bi-kvil。