Large language models encode a vast amount of semantic knowledge and possess remarkable understanding and reasoning capabilities. Previous research has explored how to ground language models in robotic tasks to ensure that the sequences generated by the language model are both logically correct and practically executable. However, low-level execution may deviate from the high-level plan due to environmental perturbations or imperfect controller design. In this paper, we propose DoReMi, a novel language model grounding framework that enables immediate Detection and Recovery from Misalignments between plan and execution. Specifically, during low-level skill execution, we use a vision question answering (VQA) model to regularly detect plan-execution misalignments. If certain misalignment occurs, our method will call the language model to re-plan in order to recover from misalignments. Experiments on various complex tasks including robot arms and humanoid robots demonstrate that our method can lead to higher task success rates and shorter task completion times. Videos of DoReMi are available at https://sites.google.com/view/doremi-paper.
翻译:大型语言模型编码了大量语义知识,并具备卓越的理解与推理能力。先前研究探讨了如何在机器人任务中对语言模型进行 grounding,以确保语言模型生成的序列既符合逻辑又具备实际可执行性。然而,由于环境干扰或控制器设计不完美,底层执行可能偏离高层计划。本文提出 DoReMi,一种新型语言模型 grounding 框架,能够即时检测并恢复计划与执行之间的失调。具体而言,在底层技能执行过程中,我们使用视觉问答(VQA)模型定期检测计划-执行失调。若发生失调,我们的方法将调用语言模型重新规划以恢复一致性。在包含机械臂与仿人机器人的多项复杂任务实验表明,我们的方法可带来更高的任务成功率和更短的任务完成时间。DoReMi 的演示视频见 https://sites.google.com/view/doremi-paper。