Large Language Models (LLMs) have recently shown promise as high-level planners for robots when given access to a selection of low-level skills. However, it is often assumed that LLMs do not possess sufficient knowledge to be used for the low-level trajectories themselves. In this work, we address this assumption thoroughly, and investigate if an LLM (GPT-4) can directly predict a dense sequence of end-effector poses for manipulation skills, when given access to only object detection and segmentation vision models. We study how well a single task-agnostic prompt, without any in-context examples, motion primitives, or external trajectory optimisers, can perform across 26 real-world language-based tasks, such as "open the bottle cap" and "wipe the plate with the sponge", and we investigate which design choices in this prompt are the most effective. Our conclusions raise the assumed limit of LLMs for robotics, and we reveal for the first time that LLMs do indeed possess an understanding of low-level robot control sufficient for a range of common tasks, and that they can additionally detect failures and then re-plan trajectories accordingly. Videos, code, and prompts are available at: https://www.robot-learning.uk/language-models-trajectory-generators.
翻译:大型语言模型(LLMs)近期在接入多种低级技能时,展现出作为机器人高层规划器的潜力。然而,人们通常认为LLMs不具备足够知识直接用于生成低级轨迹。本研究彻底验证这一假设,探究当仅接入物体检测与分割视觉模型时,LLM(GPT-4)能否直接预测操纵技能所需的密集末端执行器位姿序列。我们研究单一任务无关提示(无任何上下文示例、运动基元或外部轨迹优化器)在26项真实世界语言任务(如"打开瓶盖"和"用海绵擦盘子")中的表现,并分析提示中哪些设计选择最为有效。研究结论突破了LLMs在机器人领域的假设上限,首次揭示LLMs确实具备足以执行常见任务的低级机器人控制理解能力,且能自主检测失败并据此重新规划轨迹。相关视频、代码及提示词详见:https://www.robot-learning.uk/language-models-trajectory-generators