This paper focuses on the area of RGB(visible)-NIR(near-infrared) cross-modality image registration, which is crucial for many downstream vision tasks to fully leverage the complementary information present in visible and infrared images. In this field, researchers face two primary challenges - the absence of a correctly-annotated benchmark with viewpoint variations for evaluating RGB-NIR cross-modality registration methods and the problem of inconsistent local features caused by the appearance discrepancy between RGB-NIR cross-modality images. To address these challenges, we first present the RGB-NIR Image Registration (RGB-NIR-IRegis) benchmark, which, for the first time, enables fair and comprehensive evaluations for the task of RGB-NIR cross-modality image registration. Evaluations of previous methods highlight the significant challenges posed by our RGB-NIR-IRegis benchmark, especially on RGB-NIR image pairs with viewpoint variations. To analyze the causes of the unsatisfying performance, we then design several metrics to reveal the toxic impact of inconsistent local features between visible and infrared images on the model performance. This further motivates us to develop a baseline method named Semantic Guidance Transformer (SGFormer), which utilizes high-level semantic guidance to mitigate the negative impact of local inconsistent features. Despite the simplicity of our motivation, extensive experimental results show the effectiveness of our method.
翻译:本文聚焦于RGB(可见光)-NIR(近红外)跨模态图像配准领域,该技术对于许多下游视觉任务充分利用可见光与红外图像中的互补信息至关重要。在该领域中,研究者面临两大主要挑战:一是缺乏具有视角变化的、标注正确的基准数据集来评估RGB-NIR跨模态配准方法;二是RGB-NIR跨模态图像间外观差异导致的局部特征不一致问题。为应对这些挑战,我们首先提出了RGB-NIR图像配准(RGB-NIR-IRegis)基准数据集,首次为RGB-NIR跨模态图像配准任务提供了公平且全面的评估体系。对现有方法的评估凸显了我们的RGB-NIR-IRegis基准带来的显著挑战,尤其是在具有视角变化的RGB-NIR图像对上。为分析性能不佳的原因,我们设计了若干度量指标,以揭示可见光与红外图像间局部特征不一致对模型性能的有害影响。这进一步促使我们开发了一种名为语义引导Transformer(SGFormer)的基线方法,该方法利用高层语义引导来缓解局部不一致特征的负面影响。尽管我们的动机简洁明了,大量实验结果证明了该方法的有效性。