Planning is an important capability of artificial agents that perform long-horizon tasks in real-world environments. In this work, we explore the use of pre-trained language models (PLMs) to reason about plan sequences from text instructions in embodied visual environments. Prior PLM based approaches for planning either assume observations are available in the form of text (e.g., provided by a captioning model), reason about plans from the instruction alone, or incorporate information about the visual environment in limited ways (such as a pre-trained affordance function). In contrast, we show that PLMs can accurately plan even when observations are directly encoded as input prompts for the PLM. We show that this simple approach outperforms prior approaches in experiments on the ALFWorld and VirtualHome benchmarks.
翻译:规划是智能体在真实世界环境中执行长期任务的重要能力。本文探索利用预训练语言模型(PLM)在具身视觉环境中从文本指令推理规划序列。先前基于PLM的规划方法要么假设以文本形式(如由图像描述模型提供)获取观测信息,仅从指令本身推理规划,要么以有限方式(如预训练可泛化功能函数)融入视觉环境信息。与此不同,我们证明即便将观测信息直接编码为PLM的输入提示,PLM仍能进行精准规划。实验表明,这种简单方法在ALFWorld和VirtualHome基准测试中优于先前方法。