Gödel agent realize recursive self-improvement: an agent inspects its own policy and traces and then modifies that policy in a tested loop. We introduce Polaris, a Gödel agent for compact models that performs policy repair via experience abstraction, turning failures into policy updates through a structured cycle of analysis, strategy formation, abstraction, and minimal code pat ch repair with conservative checks. Unlike response level self correction or parameter tuning, Polaris makes policy level changes with small, auditable patches that persist in the policy and are reused on unseen instances within each benchmark. As part of the loop, the agent engages in meta reasoning: it explains its errors, proposes concrete revisions to its own policy, and then updates the policy. To enable cumulative policy refinement, we introduce experience abstraction, which distills failures into compact, reusable strategies that transfer to unseen instances. On MGSM, DROP, GPQA, and LitBench (covering arithmetic reasoning, compositional inference, graduate-level problem solving, and creative writing evaluation), a 7-billion-parameter model equipped with Polaris achieves consistent gains over the base policy and competitive baselines.
翻译:摘要:哥德尔智能体(Gödel agent)实现递归式自我改进:智能体检查其自身策略与轨迹,并在经过验证的循环中修改该策略。我们提出 Polaris——一种面向紧凑模型的哥德尔智能体,通过经验抽象进行策略修复,将失败转化为策略更新。该过程采用结构化循环,涵盖分析、策略形成、抽象、最小化代码补丁修复及保守性检查。与响应级自我纠正或参数调优不同,Polaris 通过小型可审计补丁实现策略级更改,这些补丁持久存在于策略中,并在每个基准测试中用于未见实例。在该循环中,智能体进行元推理:解释自身错误,提出针对其策略的具体修订方案,并最终更新策略。为支持累积性策略优化,我们引入经验抽象机制,将失败案例提炼为可迁移至未见实例的紧凑可重用策略。在涵盖算术推理、组合推理、研究生级问题求解与创意写作评估的 MGSM、DROP、GPQA 及 LitBench 基准测试中,配备 Polaris 的 70 亿参数模型较基础策略及竞争基线均取得持续增益。