In this work, from a theoretical lens, we aim to understand why large language model (LLM) empowered agents are able to solve decision-making problems in the physical world. To this end, consider a hierarchical reinforcement learning (RL) model where the LLM Planner and the Actor perform high-level task planning and low-level execution, respectively. Under this model, the LLM Planner navigates a partially observable Markov decision process (POMDP) by iteratively generating language-based subgoals via prompting. Under proper assumptions on the pretraining data, we prove that the pretrained LLM Planner effectively performs Bayesian aggregated imitation learning (BAIL) through in-context learning. Additionally, we highlight the necessity for exploration beyond the subgoals derived from BAIL by proving that naively executing the subgoals returned by LLM leads to a linear regret. As a remedy, we introduce an $\epsilon$-greedy exploration strategy to BAIL, which is proven to incur sublinear regret when the pretraining error is small. Finally, we extend our theoretical framework to include scenarios where the LLM Planner serves as a world model for inferring the transition model of the environment and to multi-agent settings, enabling coordination among multiple Actors.
翻译:本文从理论视角出发,旨在探究大型语言模型(LLM)赋能的智能体为何能够解决物理世界中的决策问题。为此,我们构建了一个分层强化学习(RL)模型,其中LLM规划器与执行器分别负责高层任务规划与低层动作执行。在此模型框架下,LLM规划器通过提示迭代生成基于语言的子目标,以导航部分可观测马尔可夫决策过程(POMDP)。在预训练数据满足适当假设的前提下,我们证明预训练的LLM规划器能够通过上下文学习有效执行贝叶斯聚合模仿学习(BAIL)。此外,通过理论证明机械执行LLM生成的子目标将导致线性遗憾,我们阐明了超越BAIL衍生子目标进行探索的必要性。作为改进方案,我们为BAIL引入$\epsilon$-贪心探索策略,并证明当预训练误差较小时该策略可实现次线性遗憾。最后,我们将理论框架扩展至两种场景:LLM规划器作为推断环境转移模型的世界模型,以及多智能体场景中实现多个执行器间的协同协作。