Navigating in unseen environments is crucial for mobile robots. Enhancing them with the ability to follow instructions in natural language will further improve navigation efficiency in unseen cases. However, state-of-the-art (SOTA) vision-and-language navigation (VLN) methods are mainly evaluated in simulation, neglecting the complex and noisy real world. Directly transferring SOTA navigation policies trained in simulation to the real world is challenging due to the visual domain gap and the absence of prior knowledge about unseen environments. In this work, we propose a novel navigation framework to address the VLN task in the real world. Utilizing the powerful foundation models, the proposed framework includes four key components: (1) an LLMs-based instruction parser that converts the language instruction into a sequence of pre-defined macro-action descriptions, (2) an online visual-language mapper that builds a real-time visual-language map to maintain a spatial and semantic understanding of the unseen environment, (3) a language indexing-based localizer that grounds each macro-action description into a waypoint location on the map, and (4) a DD-PPO-based local controller that predicts the action. We evaluate the proposed pipeline on an Interbotix LoCoBot WX250 in an unseen lab environment. Without any fine-tuning, our pipeline significantly outperforms the SOTA VLN baseline in the real world.
翻译:在未见环境中导航对移动机器人至关重要。增强其遵循自然语言指令的能力将进一步提升其在未见场景中的导航效率。然而,当前最先进的视觉语言导航方法主要在仿真环境中评估,忽略了真实世界的复杂性和噪声干扰。由于视觉域差距以及缺乏对未知环境的先验知识,直接将仿真训练的最优导航策略迁移到真实世界具有挑战性。本文提出一种新型导航框架以解决真实世界的视觉语言导航任务。该框架利用强大的基础模型,包含四个关键组件:(1)基于大语言模型的指令解析器,将语言指令转化为预定义的宏动作描述序列;(2)在线视觉语言地图构建模块,实时构建视觉语言地图以维持对未知环境的空间与语义理解;(3)基于语言索引的定位器,将每个宏动作描述映射至地图中的路径点位置;(4)基于DD-PPO的局部控制器,预测具体动作。我们在未见实验室环境中基于Interbotix LoCoBot WX250平台评估了所提流水线。无需任何微调,该流水线在真实世界中显著超越当前最先进的视觉语言导航基线方法。