To fully leverage the capabilities of mobile manipulation robots, it is imperative that they are able to autonomously execute long-horizon tasks in large unexplored environments. While large language models (LLMs) have shown emergent reasoning skills on arbitrary tasks, existing work primarily concentrates on explored environments, typically focusing on either navigation or manipulation tasks in isolation. In this work, we propose MoMa-LLM, a novel approach that grounds language models within structured representations derived from open-vocabulary scene graphs, dynamically updated as the environment is explored. We tightly interleave these representations with an object-centric action space. The resulting approach is zero-shot, open-vocabulary, and readily extendable to a spectrum of mobile manipulation and household robotic tasks. We demonstrate the effectiveness of MoMa-LLM in a novel semantic interactive search task in large realistic indoor environments. In extensive experiments in both simulation and the real world, we show substantially improved search efficiency compared to conventional baselines and state-of-the-art approaches, as well as its applicability to more abstract tasks. We make the code publicly available at http://moma-llm.cs.uni-freiburg.de.
翻译:为充分释放移动操作机器人的潜力,必须使其能够在未探索的大型环境中自主执行长时域任务。尽管大语言模型在任意任务中展现出涌现推理能力,但现有研究主要集中于已探索环境,且通常单独聚焦导航或操作任务。本文提出MoMa-LLM,这是一种创新方法,将语言模型嵌入基于开放词汇场景图的结构化表征中,并在环境探索过程中动态更新该表征。我们将这些表征与以物体为中心的动作空间紧密交织。最终方法具有零样本、开放词汇特性,并可直接扩展至一系列移动操作与家庭机器人任务。我们在大型逼真室内环境中设计的新型语义交互式搜索任务上验证了MoMa-LLM的有效性。在仿真与真实场景的大量实验中,相较于传统基线及最先进方法,我们展示了显著提升的搜索效率及其在更抽象任务中的适用性。代码已开源至http://moma-llm.cs.uni-freiburg.de。