Semantic 2D maps are commonly used by humans and machines for navigation purposes, whether it's walking or driving. However, these maps have limitations: they lack detail, often contain inaccuracies, and are difficult to create and maintain, especially in an automated fashion. Can we use raw imagery to automatically create better maps that can be easily interpreted by both humans and machines? We introduce SNAP, a deep network that learns rich neural 2D maps from ground-level and overhead images. We train our model to align neural maps estimated from different inputs, supervised only with camera poses over tens of millions of StreetView images. SNAP can resolve the location of challenging image queries beyond the reach of traditional methods, outperforming the state of the art in localization by a large margin. Moreover, our neural maps encode not only geometry and appearance but also high-level semantics, discovered without explicit supervision. This enables effective pre-training for data-efficient semantic scene understanding, with the potential to unlock cost-efficient creation of more detailed maps.
翻译:语义二维地图常被人类和机器用于导航任务(无论是步行还是驾驶)。然而,这些地图存在局限性:缺乏细节、常含不准确之处,且创建与维护困难(尤其在自动化场景下)。我们能否利用原始影像自动构建更优质的地图,使其同时被人类和机器轻松解读?本文提出SNAP——一种从地面及空中图像中学习丰富神经二维地图的深度网络。我们训练模型对齐从不同输入估计的神经地图,仅通过数千万张街景图像的相机位姿进行监督。SNAP能够解析传统方法难以处理的挑战性图像查询定位问题,以极大优势超越现有最优定位方法。此外,我们的神经地图不仅编码几何特征与视觉外观,还隐式发现高层语义信息(无需显式监督)。这为数据高效型语义场景理解提供了有效预训练手段,有望解锁更具成本效益的高精细地图创建方式。