Humans localize themselves efficiently in known environments by first recognizing landmarks defined on certain objects and their spatial relationships, and then verifying the location by aligning detailed structures of recognized objects with those in the memory. Inspired by this, we propose the place recognition anywhere model (PRAM) to perform visual localization as efficiently as humans do. PRAM consists of two main components - recognition and registration. In detail, first of all, a self-supervised map-centric landmark definition strategy is adopted, making places in either indoor or outdoor scenes act as unique landmarks. Then, sparse keypoints extracted from images, are utilized as the input to a transformer-based deep neural network for landmark recognition; these keypoints enable PRAM to recognize hundreds of landmarks with high time and memory efficiency. Keypoints along with recognized landmark labels are further used for registration between query images and the 3D landmark map. Different from previous hierarchical methods, PRAM discards global and local descriptors, and reduces over 90% storage. Since PRAM utilizes recognition and landmark-wise verification to replace global reference search and exhaustive matching respectively, it runs 2.4 times faster than prior state-of-the-art approaches. Moreover, PRAM opens new directions for visual localization including multi-modality localization, map-centric feature learning, and hierarchical scene coordinate regression.
翻译:人类通过先识别由特定物体及其空间关系定义的地标,再通过将识别物体的详细结构与记忆中的结构对齐来验证位置,从而在已知环境中高效地进行定位。受此启发,我们提出任意场所识别模型(PRAM),以像人类一样高效地执行视觉定位。PRAM由两个主要组件构成——识别与配准。具体而言,首先采用一种自监督的以地图为中心的地标定义策略,使室内或室外场景中的位置均可作为独特地标。然后,利用从图像中提取的稀疏关键点作为基于Transformer的深度神经网络的输入进行地标识别;这些关键点使PRAM能够以高时间和内存效率识别数百个地标。关键点连同识别到的地标标签进一步用于查询图像与三维地标地图之间的配准。与以往的分层方法不同,PRAM摒弃了全局与局部描述子,并减少了90%以上的存储开销。由于PRAM分别用识别和逐地标验证取代了全局参考搜索与穷举匹配,其运行速度比现有最先进方法快2.4倍。此外,PRAM为视觉定位开辟了新方向,包括多模态定位、以地图为中心的特征学习以及分层场景坐标回归。