Enabling robots to follow complex natural language instructions is an important yet challenging problem. People want to flexibly express constraints, refer to arbitrary landmarks and verify behavior when instructing robots. Conversely, robots must disambiguate human instructions into specifications and ground instruction referents in the real world. We propose Language Instruction grounding for Motion Planning (LIMP), a system that leverages foundation models and temporal logics to generate instruction-conditioned semantic maps that enable robots to verifiably follow expressive and long-horizon instructions with open vocabulary referents and complex spatiotemporal constraints. In contrast to prior methods for using foundation models in robot task execution, LIMP constructs an explainable instruction representation that reveals the robot's alignment with an instructor's intended motives and affords the synthesis of robot behaviors that are correct-by-construction. We demonstrate LIMP in three real-world environments, across a set of 35 complex spatiotemporal instructions, showing the generality of our approach and the ease of deployment in novel unstructured domains. In our experiments, LIMP can spatially ground open-vocabulary referents and synthesize constraint-satisfying plans in 90% of object-goal navigation and 71% of mobile manipulation instructions. See supplementary videos at https://robotlimp.github.io
翻译:让机器人遵循复杂的自然语言指令是一项重要且具有挑战性的问题。人类希望灵活表达约束条件、引用任意地标并验证机器人行为;而机器人必须将人类指令消歧为规范形式,并将指令指代物映射到现实世界。我们提出面向运动规划的语言指令接地系统(LIMP),该系统利用基础模型和时序逻辑生成指令条件语义地图,使机器人能够以开放词汇指代物和复杂时空约束,可验证地执行具有表现力且长时跨的指令。与现有在机器人任务执行中使用基础模型的方法不同,LIMP构建了可解释的指令表征——既揭示机器人对操作者意图的对齐程度,又支持生成"构造即正确"的机器人行为。我们在三个真实世界环境中,针对35条复杂时空指令展示了LIMP的通用性及其在新型非结构化场景中的易部署性。实验表明,LIMP能实现开放词汇指代物的空间接地,并在90%的物体目标导航任务和71%的移动操作指令中生成满足约束的规划方案。补充视频见https://robotlimp.github.io