Large language models (LLMs) demonstrate their promise in tackling complicated practical challenges by combining action-based policies with chain of thought (CoT) reasoning. Having high-quality prompts on hand, however, is vital to the framework's effectiveness. Currently, these prompts are handcrafted utilizing extensive human labor, resulting in CoT policies that frequently fail to generalize. Human intervention is also required in order to develop grounding functions that ensure low-level controllers appropriately process CoT reasoning. In this paper, we take the first step towards a fully integrated end-to-end framework for task-solving in real settings employing complicated reasoning. To that purpose, we offer a new leader-follower bilevel framework capable of learning to ask relevant questions (prompts) and subsequently undertaking reasoning to guide the learning of actions to be performed in an environment. A good prompt should make introspective revisions based on historical findings, leading the CoT to consider the anticipated goals. A prompt-generator policy has its own aim in our system, allowing it to adapt to the action policy and automatically root the CoT process towards outputs that lead to decisive, high-performing actions. Meanwhile, the action policy is learning how to use the CoT outputs to take specific actions. Our empirical data reveal that our system outperforms leading methods in agent learning benchmarks such as Overcooked and FourRoom.
翻译:大型语言模型通过将基于行动的策略与思维链推理相结合,展现出应对复杂实际挑战的潜力。然而,拥有高质量的提示对这一框架的效能至关重要。当前,这些提示依赖大量人工精心设计,导致思维链策略常难以泛化。此外,为确保低级控制器正确处理思维链推理,还需人工开发接地函数。本文首次迈出构建完全集成端到端框架以解决实际环境中复杂推理任务的第一步。为此,我们提出一种新的领导者-跟随者双层框架,能够学习提出相关问题(提示),随后进行推理以指导环境中待执行动作的学习。一个良好的提示应基于历史发现进行内省修正,引导思维链聚焦预期目标。在我们的系统中,提示生成策略拥有自身目标,使其能适应行动策略,并自动将思维链过程导向产生决定性、高性能行动的输出。与此同时,行动策略则学习如何利用思维链输出执行具体动作。实验数据表明,我们的系统在Overcooked和FourRoom等智能体学习基准测试中优于主流方法。