This paper introduces a novel deep-learning approach for human-to-robot motion retargeting, enabling robots to mimic human poses accurately. Contrary to prior deep-learning-based works, our method does not require paired human-to-robot data, which facilitates its translation to new robots. First, we construct a shared latent space between humans and robots via adaptive contrastive learning that takes advantage of a proposed cross-domain similarity metric between the human and robot poses. Additionally, we propose a consistency term to build a common latent space that captures the similarity of the poses with precision while allowing direct robot motion control from the latent space. For instance, we can generate in-between motion through simple linear interpolation between two projected human poses. We conduct a comprehensive evaluation of robot control from diverse modalities (i.e., texts, RGB videos, and key poses), which facilitates robot control for non-expert users. Our model outperforms existing works regarding human-to-robot retargeting in terms of efficiency and precision. Finally, we implemented our method in a real robot with self-collision avoidance through a whole-body controller to showcase the effectiveness of our approach. More information on our website https://evm7.github.io/UnsH2R/
翻译:本文提出了一种新颖的深度学习方法,用于实现人-机器人运动重定向,使机器人能够精准模仿人类姿态。与以往基于深度学习的工作不同,本方法无需配对的人-机器人数据,从而便于将其迁移至新型机器人。首先,我们通过自适应对比学习构建人与机器人之间的共享潜在空间,该学习过程利用了提出的人与机器人姿态间的跨域相似性度量。此外,我们引入一致性项来构建一个共同潜在空间,该空间既能精确捕捉姿态的相似性,又允许从潜在空间直接控制机器人运动。例如,通过两个投影人体姿态间的简单线性插值即可生成中间运动。我们从多种模态(即文本、RGB视频和关键姿态)对机器人控制进行全面评估,从而便利非专业用户的操作。在效率与精度方面,我们的模型在人-机器人重定向任务上优于现有工作。最后,我们在真实机器人上通过全身控制器实现了本方法,并具备自碰撞规避能力,以展示其有效性。更多信息请访问我们的网站 https://evm7.github.io/UnsH2R/