It is a long-lasting goal to design an embodied system that can solve long-horizon open-world tasks in human-like ways. However, existing approaches usually struggle with compound difficulties caused by the logic-aware decomposition and context-aware execution of these tasks. To this end, we introduce MP5, an open-ended multimodal embodied system built upon the challenging Minecraft simulator, which can decompose feasible sub-objectives, design sophisticated situation-aware plans, and perform embodied action control, with frequent communication with a goal-conditioned active perception scheme. Specifically, MP5 is developed on top of recent advances in Multimodal Large Language Models (MLLMs), and the system is modulated into functional modules that can be scheduled and collaborated to ultimately solve pre-defined context- and process-dependent tasks. Extensive experiments prove that MP5 can achieve a 22% success rate on difficult process-dependent tasks and a 91% success rate on tasks that heavily depend on the context. Moreover, MP5 exhibits a remarkable ability to address many open-ended tasks that are entirely novel.
翻译:设计一个能够以类人方式解决长时序开放式世界任务的具身系统是一个长期目标。然而,现有方法通常难以应对由这些任务的逻辑感知分解和上下文感知执行所导致的复合性困难。为此,我们提出了MP5——一个基于具有挑战性的Minecraft模拟器构建的开放式多模态具身系统,该系统通过频繁与目标条件主动感知方案通信,能够分解可行的子目标、设计复杂的情境感知规划并执行具身动作控制。具体而言,MP5基于多模态大语言模型(MLLMs)的最新进展开发,系统被模块化为可调度与协作的功能模块,最终解决预定义的上下文依赖与过程依赖任务。大量实验证明,MP5在困难的过程依赖任务上可实现22%的成功率,在高度依赖上下文的任务上可实现91%的成功率。此外,MP5在解决完全新颖的开放式任务方面展现出卓越能力。