We study the image-based geolocalization problem, aiming to localize ground-view query images on cartographic maps. Current methods often utilize cross-view localization techniques to match ground-view query images with 2D maps. However, the performance of these methods is unsatisfactory due to significant cross-view appearance differences. In this paper, we lift cross-view matching to a 2.5D space, where heights of structures (e.g., trees and buildings) provide geometric information to guide the cross-view matching. We propose a new approach to learning representative embeddings from multi-modal data. Specifically, we establish a projection relationship between 2.5D space and 2D aerial-view space. The projection is further used to combine multi-modal features from the 2.5D and 2D maps using an effective pixel-to-point fusion method. By encoding crucial geometric cues, our method learns discriminative location embeddings for matching panoramic images and maps. Additionally, we construct the first large-scale ground-to-2.5D map geolocalization dataset to validate our method and facilitate future research. Both single-image based and route based localization experiments are conducted to test our method. Extensive experiments demonstrate that the proposed method achieves significantly higher localization accuracy and faster convergence than previous 2D map-based approaches.
翻译:我们研究基于图像的地理定位问题,旨在将地面视角的查询图像定位到地图上。当前方法常利用跨视角定位技术,将地面视角的查询图像与二维地图进行匹配。然而,由于显著的跨视角外观差异,这些方法的性能并不理想。本文中,我们将跨视角匹配提升至2.5D空间,其中结构(如树木和建筑物)的高度提供了几何信息以指导跨视角匹配。我们提出了一种新方法,从多模态数据中学习具有代表性的嵌入表示。具体而言,我们建立了2.5D空间与二维航拍视角空间之间的投影关系。该投影进一步用于通过有效的像素到点融合方法,结合来自2.5D和二维地图的多模态特征。通过编码关键的几何线索,我们的方法学习了具有判别性的位置嵌入,用于匹配全景图像和地图。此外,我们构建了首个大规模地面到2.5D地图地理定位数据集,以验证我们的方法并促进未来研究。我们进行了基于单张图像和基于路线的定位实验来测试我们的方法。大量实验表明,与之前基于二维地图的方法相比,所提方法在定位精度和收敛速度上均有显著提升。