Learning generalizable visual dynamic representation across different embodied environments is crucial for real-world robotic manipulation. As the scale and diversity of robot demonstration data are limited, recent works have turned to large-scale pre-training using human data. However, the morphological differences between humans and robots introduce a significant human-robot domain discrepancy, challenging the generalization of these human-data pre-trained models to downstream manipulation tasks. To address this, we propose a novel adaptation paradigm that utilizes readily available paired human-robot video data to bridge the discrepancy. Following this paradigm, our method exploits a human-robot contrastive alignment loss to align the semantics of human and robot videos, adapting pre-trained models to the robotic domain in a parameter-efficient manner. The experiments demonstrate significant improvements on 25 tasks across three different benchmarks, where the single-task, language-conditioned multi-task settings are covered, and two different pre-trained models are evaluated. On the large RLBench benchmark, our adaptation method achieves an average improvement of $8.9\%$ in success rate over the pre-trained R3M model across multiple tasks. We will release the code and models upon acceptance.
翻译:学习跨不同具身环境的通用视觉动态表征对于现实世界的机器人操作至关重要。由于机器人示范数据的规模和多样性有限,近期研究转向利用人类数据进行大规模预训练。然而,人与机器人之间的形态差异引入了显著的人-机器人领域差异,这挑战了这些基于人类数据预训练的模型在下游操作任务中的泛化能力。为解决此问题,我们提出一种新颖的适应范式,利用易于获取的配对人-机器人视频数据来弥合该差异。遵循此范式,我们的方法采用人-机器人对比对齐损失来对齐人类与机器人视频的语义,以参数高效的方式使预训练模型适应机器人领域。实验在涵盖三个不同基准的25项任务上展示了显著改进,其中覆盖了单任务和语言条件多任务设置,并评估了两种不同的预训练模型。在大型RLBench基准测试中,我们的适应方法在多项任务上相比预训练的R3M模型平均成功率提升了$8.9\%$。我们将在论文被接受后公开代码和模型。