This paper presents the Embedding Pose Graph (EPG), an innovative method that combines the strengths of foundation models with a simple 3D representation suitable for robotics applications. Addressing the need for efficient spatial understanding in robotics, EPG provides a compact yet powerful approach by attaching foundation model features to the nodes of a pose graph. Unlike traditional methods that rely on bulky data formats like voxel grids or point clouds, EPG is lightweight and scalable. It facilitates a range of robotic tasks, including open-vocabulary querying, disambiguation, image-based querying, language-directed navigation, and re-localization in 3D environments. We showcase the effectiveness of EPG in handling these tasks, demonstrating its capacity to improve how robots interact with and navigate through complex spaces. Through both qualitative and quantitative assessments, we illustrate EPG's strong performance and its ability to outperform existing methods in re-localization. Our work introduces a crucial step forward in enabling robots to efficiently understand and operate within large-scale 3D spaces.
翻译:本文提出嵌入位姿图(EPG),这是一种融合基础模型优势与适用于机器人应用的简化三维表示的创新方法。针对机器人对高效空间感知的需求,EPG通过将基础模型特征附加到位姿图节点上,提供了一种紧凑而强大的解决方案。不同于依赖体素网格或点云等庞大数据格式的传统方法,EPG轻量级且可扩展,支持多项机器人任务,包括开放词汇查询、歧义消解、基于图像的查询、语言引导导航及三维环境中的重定位。我们展示了EPG在处理这些任务中的有效性,论证了其提升机器人在复杂空间中交互与导航能力。通过定性与定量评估,我们证实了EPG的卓越性能及其在重定位任务中超越现有方法的能力。本研究为促进机器人在大规模三维空间中高效理解与操作迈出了关键一步。