This paper focuses on the sim-to-real issue of RGB-D grasp detection and formulates it as a domain adaptation problem. In this case, we present a global-to-local method to address hybrid domain gaps in RGB and depth data and insufficient multi-modal feature alignment. First, a self-supervised rotation pre-training strategy is adopted to deliver robust initialization for RGB and depth networks. We then propose a global-to-local alignment pipeline with individual global domain classifiers for scene features of RGB and depth images as well as a local one specifically working for grasp features in the two modalities. In particular, we propose a grasp prototype adaptation module, which aims to facilitate fine-grained local feature alignment by dynamically updating and matching the grasp prototypes from the simulation and real-world scenarios throughout the training process. Due to such designs, the proposed method substantially reduces the domain shift and thus leads to consistent performance improvements. Extensive experiments are conducted on the GraspNet-Planar benchmark and physical environment, and superior results are achieved which demonstrate the effectiveness of our method.
翻译:本文聚焦于RGB-D抓取检测的仿真到真实场景迁移问题,并将其形式化为域适应任务。针对RGB与深度数据中存在的混合域差异以及多模态特征对齐不足的问题,我们提出了一种全局到局部的处理方法。首先,采用自监督旋转预训练策略为RGB和深度网络提供鲁棒的初始化参数。继而构建全局到局部对齐框架:通过为RGB与深度图像的场景特征分别设置独立的全局域分类器,同时为两种模态的抓取特征设计专用的局部域分类器。特别地,我们提出抓取原型适配模块,该模块通过在整个训练过程中动态更新并匹配仿真域与真实域的抓取原型,实现细粒度局部特征对齐。基于上述设计,所提方法显著降低了域偏移,从而持续提升性能。在GraspNet-Planar基准测试集与物理环境中的大量实验表明,该方法取得了优越性能,充分验证了其有效性。