Cross-view geo-localization aims to estimate the location of a query ground image by matching it to a reference geo-tagged aerial images database. As an extremely challenging task, its difficulties root in the drastic view changes and different capturing time between two views. Despite these difficulties, recent works achieve outstanding progress on cross-view geo-localization benchmarks. However, existing methods still suffer from poor performance on the cross-area benchmarks, in which the training and testing data are captured from two different regions. We attribute this deficiency to the lack of ability to extract the spatial configuration of visual feature layouts and models' overfitting on low-level details from the training set. In this paper, we propose GeoDTR which explicitly disentangles geometric information from raw features and learns the spatial correlations among visual features from aerial and ground pairs with a novel geometric layout extractor module. This module generates a set of geometric layout descriptors, modulating the raw features and producing high-quality latent representations. In addition, we elaborate on two categories of data augmentations, (i) Layout simulation, which varies the spatial configuration while keeping the low-level details intact. (ii) Semantic augmentation, which alters the low-level details and encourages the model to capture spatial configurations. These augmentations help to improve the performance of the cross-view geo-localization models, especially on the cross-area benchmarks. Moreover, we propose a counterfactual-based learning process to benefit the geometric layout extractor in exploring spatial information. Extensive experiments show that GeoDTR not only achieves state-of-the-art results but also significantly boosts the performance on same-area and cross-area benchmarks.
翻译:跨视角地理定位旨在通过将查询地面图像与参考地理标记的航拍图像数据库进行匹配,来估计其位置。作为一项极具挑战的任务,其难点源于两种视角间的剧烈视角变化和不同的拍摄时间。尽管存在这些困难,近期工作仍在跨视角地理定位基准测试中取得了显著进展。然而,现有方法在跨区域基准测试中仍表现不佳——这些基准的训练和测试数据来自两个不同区域。我们将此缺陷归因于模型缺乏提取视觉特征布局空间配置的能力,以及过度拟合训练集中的低级细节。本文提出GeoDTR,该方法明确从原始特征中解耦几何信息,并通过新颖的几何布局提取器模块学习航拍与地面图像对中视觉特征间的空间关联。该模块生成一组几何布局描述符,对原始特征进行调制,从而产生高质量的潜在表征。此外,我们精心设计了两种数据增强方法:(i)布局模拟,在保持低级细节不变的同时改变空间配置;(ii)语义增强,通过改变低级细节促使模型捕捉空间配置。这些增强方法有助于提升跨视角地理定位模型的性能,尤其在跨区域基准测试中。进一步地,我们还提出基于反事实的学习过程,以帮助几何布局提取器探索空间信息。大量实验表明,GeoDTR不仅取得了最新最优结果,还显著提升了在同区域和跨区域基准测试中的性能。