Visual place recognition is essential for vision-based robot localization and SLAM. Despite the tremendous progress made in recent years, place recognition in changing environments remains challenging. A promising approach to cope with appearance variations is to leverage high-level semantic features like objects or place categories. In this paper, we propose FM-Loc which is a novel image-based localization approach based on Foundation Models that uses the Large Language Model GPT-3 in combination with the Visual-Language Model CLIP to construct a semantic image descriptor that is robust to severe changes in scene geometry and camera viewpoint. We deploy CLIP to detect objects in an image, GPT-3 to suggest potential room labels based on the detected objects, and CLIP again to propose the most likely location label. The object labels and the scene label constitute an image descriptor that we use to calculate a similarity score between the query and database images. We validate our approach on real-world data that exhibit significant changes in camera viewpoints and object placement between the database and query trajectories. The experimental results demonstrate that our method is applicable to a wide range of indoor scenarios without the need for training or fine-tuning.
翻译:视觉位置识别对于基于视觉的机器人定位和SLAM至关重要。尽管近年来取得了巨大进展,但在变化环境中的位置识别仍然具有挑战性。应对外观变化的一种有前景的方法是利用高级语义特征,如物体或场所类别。本文提出FM-Loc,这是一种基于基础模型的创新图像定位方法,该方法结合大型语言模型GPT-3与视觉-语言模型CLIP,构建对场景几何结构和相机视角剧烈变化具有鲁棒性的语义图像描述符。我们部署CLIP检测图像中的物体,利用GPT-3基于检测到的物体建议可能的房间标签,再通过CLIP提出最可能的位置标签。物体标签与场景标签共同构成图像描述符,用于计算查询图像与数据库图像之间的相似度分数。我们在呈现数据库轨迹与查询轨迹间相机视角和物体布局显著变化的真实场景数据上验证了该方法。实验结果表明,该方法无需训练或微调即可适用于广泛的室内场景。