We present LGX, a novel algorithm for Object Goal Navigation in a "language-driven, zero-shot manner", where an embodied agent navigates to an arbitrarily described target object in a previously unexplored environment. Our approach leverages the capabilities of Large Language Models (LLMs) for making navigational decisions by mapping the LLMs implicit knowledge about the semantic context of the environment into sequential inputs for robot motion planning. Simultaneously, we also conduct generalized target object detection using a pre-trained Vision-Language grounding model. We achieve state-of-the-art zero-shot object navigation results on RoboTHOR with a success rate (SR) improvement of over 27% over the current baseline of the OWL-ViT CLIP on Wheels (OWL CoW). Furthermore, we study the usage of LLMs for robot navigation and present an analysis of the various semantic factors affecting model output. Finally, we showcase the benefits of our approach via real-world experiments that indicate the superior performance of LGX when navigating to and detecting visually unique objects.
翻译:我们提出LGX,一种以“语言驱动、零样本方式”实现物体目标导航的新算法,其中具身智能体能在未探索环境中导航至任意描述的指定目标物体。该方法通过将大语言模型(LLMs)对环境语义上下文的隐式知识映射为机器人运动规划的序列输入,充分发挥LLMs在导航决策中的能力。同时,我们还利用预训练的视觉-语言对齐模型进行泛化目标物体检测。我们在RoboTHOR平台上实现了零样本物体导航的最新成果,成功率(SR)较现有OWL-ViT CLIP on Wheels(OWL CoW)基线提升超过27%。此外,我们研究了LLMs在机器人导航中的应用,分析了影响模型输出的多种语义因素。最后,通过真实世界实验验证了我们的方法在导航至并检测视觉独特物体时的优越性能,凸显了LGX的优势。