Embodied AI focuses on the study and development of intelligent systems that possess a physical or virtual embodiment (i.e. robots) and are able to dynamically interact with their environment. Memory and control are the two essential parts of an embodied system and usually require separate frameworks to model each of them. In this paper, we propose a novel and generalizable framework called LLM-Brain: using Large-scale Language Model as a robotic brain to unify egocentric memory and control. The LLM-Brain framework integrates multiple multimodal language models for robotic tasks, utilizing a zero-shot learning approach. All components within LLM-Brain communicate using natural language in closed-loop multi-round dialogues that encompass perception, planning, control, and memory. The core of the system is an embodied LLM to maintain egocentric memory and control the robot. We demonstrate LLM-Brain by examining two downstream tasks: active exploration and embodied question answering. The active exploration tasks require the robot to extensively explore an unknown environment within a limited number of actions. Meanwhile, the embodied question answering tasks necessitate that the robot answers questions based on observations acquired during prior explorations.
翻译:具身人工智能致力于研究并开发具有物理或虚拟实体(即机器人)且能动态与环境交互的智能系统。记忆与控制是具身系统的两个核心组成部分,通常需要分别建模。本文提出一种新颖且具备通用性的框架——LLM-Brain:将大规模语言模型作为机器人脑,以统一自我中心记忆与控制。该框架通过零样本学习方法,整合多个多模态语言模型以完成机器人任务。其所有组件在闭合循环的多轮对话中通过自然语言进行通信,涵盖感知、规划、控制与记忆。系统的核心是一个具身大语言模型,用于维护自我中心记忆并控制机器人。我们通过主动探索与具身问答两个下游任务验证LLM-Brain的性能:主动探索任务要求机器人在有限动作次数内充分探索未知环境;而具身问答任务则要求机器人基于先前探索中获取的观测信息回答提问。