Geolocation is a fundamental component of route planning and navigation for unmanned vehicles, but GNSS-based geolocation fails under denial-of-service conditions. Cross-view geo-localization (CVGL), which aims to estimate the geographical location of the ground-level camera by matching against enormous geo-tagged aerial (\emph{e.g.}, satellite) images, has received lots of attention but remains extremely challenging due to the drastic appearance differences across aerial-ground views. In existing methods, global representations of different views are extracted primarily using Siamese-like architectures, but their interactive benefits are seldom taken into account. In this paper, we present a novel approach using cross-view knowledge generative techniques in combination with transformers, namely mutual generative transformer learning (MGTL), for CVGL. Specifically, by taking the initial representations produced by the backbone network, MGTL develops two separate generative sub-modules -- one for aerial-aware knowledge generation from ground-view semantics and vice versa -- and fully exploits the entirely mutual benefits through the attention mechanism. Moreover, to better capture the co-visual relationships between aerial and ground views, we introduce a cascaded attention masking algorithm to further boost accuracy. Extensive experiments on challenging public benchmarks, \emph{i.e.}, {CVACT} and {CVUSA}, demonstrate the effectiveness of the proposed method which sets new records compared with the existing state-of-the-art models.
翻译:地理定位是无人驾驶车辆路径规划与导航的基础,但基于GNSS的定位在拒绝服务条件下会失效。跨视角地理定位(CVGL)旨在通过将地面级图像与大量带有地理标签的航拍(如卫星)图像进行匹配来估计其地理位置,该方法虽备受关注,但由于航拍-地面视角间存在显著的视觉差异而极具挑战性。现有方法主要通过Siamese类架构提取不同视角的全局表示,但很少考虑其交互优势。本文提出一种融合跨视角知识生成技术与Transformer的新方法——即互生成式Transformer学习(MGTL)——用于CVGL。具体而言,MGTL利用骨干网络生成的初始表示,构建两个独立的生成子模块(分别实现基于地面语义的航拍感知知识生成及其逆向过程),并通过注意力机制充分挖掘跨视角的完全互惠优势。此外,为更好捕捉航拍与地面视角间的共视关系,我们引入级联注意力掩蔽算法以进一步提升精度。在CVACT与CVUSA等挑战性公共基准上的大量实验表明,所提方法有效超越了现有最优模型,创下新的性能记录。