Self-driving vehicles rely on urban street maps for autonomous navigation. In this paper, we introduce Pix2Map, a method for inferring urban street map topology directly from ego-view images, as needed to continually update and expand existing maps. This is a challenging task, as we need to infer a complex urban road topology directly from raw image data. The main insight of this paper is that this problem can be posed as cross-modal retrieval by learning a joint, cross-modal embedding space for images and existing maps, represented as discrete graphs that encode the topological layout of the visual surroundings. We conduct our experimental evaluation using the Argoverse dataset and show that it is indeed possible to accurately retrieve street maps corresponding to both seen and unseen roads solely from image data. Moreover, we show that our retrieved maps can be used to update or expand existing maps and even show proof-of-concept results for visual localization and image retrieval from spatial graphs.
翻译:自动驾驶车辆依赖城市街道地图实现自主导航。本文提出Pix2Map方法,旨在直接从自车视角图像推断城市街道地图拓扑结构,以满足持续更新和扩展现有地图的需求。这是一项具有挑战性的任务,因为我们需要直接从原始图像数据中推断复杂的城市道路拓扑。本文的核心见解在于:通过学习图像与现有地图(以编码视觉环境拓扑布局的离散图表示)的联合跨模态嵌入空间,可将该问题转化为跨模态检索任务。我们利用Argoverse数据集进行实验评估,结果表明,仅凭图像数据即可准确检索对应可见及未见道路的街道地图。此外,我们证明检索得到的地图可用于更新或扩展现有地图,并展示了基于空间图进行视觉定位与图像检索的概念验证结果。