Matching cross-modality features between images and point clouds is a fundamental problem for image-to-point cloud registration. However, due to the modality difference between images and points, it is difficult to learn robust and discriminative cross-modality features by existing metric learning methods for feature matching. Instead of applying metric learning on cross-modality data, we propose to unify the modality between images and point clouds by pretrained large-scale models first, and then establish robust correspondence within the same modality. We show that the intermediate features, called diffusion features, extracted by depth-to-image diffusion models are semantically consistent between images and point clouds, which enables the building of coarse but robust cross-modality correspondences. We further extract geometric features on depth maps produced by the monocular depth estimator. By matching such geometric features, we significantly improve the accuracy of the coarse correspondences produced by diffusion features. Extensive experiments demonstrate that without any task-specific training, direct utilization of both features produces accurate image-to-point cloud registration. On three public indoor and outdoor benchmarks, the proposed method averagely achieves a 20.6 percent improvement in Inlier Ratio, a three-fold higher Inlier Number, and a 48.6 percent improvement in Registration Recall than existing state-of-the-arts.
翻译:图像与点云之间的跨模态特征匹配是图像-点云配准中的核心问题。然而,由于图像与点云间的模态差异,现有基于度量学习的特征匹配方法难以学习鲁棒且具有判别性的跨模态特征。本文提出先通过预训练大规模模型统一图像与点云的模态,再在同一模态内建立鲁棒对应关系,而非对跨模态数据直接应用度量学习。我们发现,由深度-图像扩散模型提取的中间特征(称为扩散特征)在图像与点云间具有语义一致性,能够建立粗粒度但鲁棒的跨模态对应关系。进一步,我们提取单目深度估计器生成的深度图上的几何特征,通过匹配此类几何特征,可显著提升扩散特征所产生粗对应关系的精度。大量实验表明,无需任何任务特定训练,直接联合利用这两类特征即可实现精确的图像-点云配准。在三个公开室内外基准测试中,所提方法相比现有最优方法,内点率平均提升20.6%,内点数量提升3倍,配准召回率提升48.6%。