We present ReflexGrad, a dual-process architecture for within-episode failure recovery in LLM agents without demonstrations. When agents commit to a wrong approach early and exhaust the step budget, the post-failure trajectory contains the information to escape -- but no published architecture acts on it within a single episode. ReflexGrad routes between a fast process (TextGrad-style continuous refinement every $k{=}3$ steps) and a slow process (Reflexion-style causal diagnosis when $m{=}5$ consecutive low-progress scores fire a routing gate). A deterministic priority merge keeps the natural-language policy coherent, and each slow activation emits three observable artifacts: a reproducible trigger, a causal diagnostic, and a verified fix. On ALFWorld 134 tasks, $n{=}10$ seeds, no demonstrations, ReflexGrad lifts Qwen-3-8B from $35.1\%$ to $75.4\%$ ($+40.3$pp), beating compute-matched 1-shot LATS by $+2.7$pp ($p{\approx}0.01$), ToT by $+5.7$pp ($p{<}10^{-4}$), and Self-Refine by $+6.7$pp ($p{<}10^{-5}$); on GPT-5 the lift is $46.3{\to}88.1\%$ ($+41.8$pp). The $1.5$pp cross-model difference is within seed noise ($p{\approx}0.13$), suggesting that the routing mechanism, rather than model scale, is the primary source of the gain. Code, prompts, per-seed logs, and sensitivity sweeps are released.


翻译:我们提出ReflexGrad——一种无需示范即可实现LLM智能体片段内故障恢复的双过程架构。当智能体早期错误决策导致步骤预算耗尽时,故障后轨迹中实际蕴含逃逸信息,但现有架构均未能在单一片段内对此加以利用。ReflexGrad在快速过程(每$k{=}3$步执行TextGrad风格连续优化)与慢速过程(当$m{=}5$个连续低进度分数触发路由门时执行Reflexion风格因果诊断)之间切换路由。确定性优先级合并机制确保自然语言策略的一致性,每次慢速激活生成三个可观测产物:可复现触发器、因果诊断结果与已验证修复方案。在ALFWorld 134个任务的$n{=}10$种子实验中(零示范条件下),ReflexGrad将Qwen-3-8B从35.1%提升至75.4%(提升40.3个百分点),以$+2.7$个百分点($p{\approx}0.01$)超越计算量匹配的单次LATS,以$+5.7$个百分点($p{<}10^{-4}$)超越ToT,以$+6.7$个百分点($p{<}10^{-5}$)超越Self-Refine;在GPT-5上升幅为46.3%→88.1%(提升41.8个百分点)。1.5个百分点的跨模型差异在种子噪声范围内($p{\approx}0.13$),表明增益主要来源于路由机制而非模型规模。代码、提示词、逐种子日志及敏感性扫描已开源。

0
下载
关闭预览

相关内容

综述 | Always-On Agents:LLM智能体持久状态与治理
专知会员服务
22+阅读 · 6月30日
AgentOps综述:智能体系统运维框架
专知会员服务
24+阅读 · 6月4日
分布式并行架构Ray介绍
CreateAMind
10+阅读 · 2019年8月9日
基于车路协同的群体智能协同
智能交通技术
10+阅读 · 2019年1月23日
三次简化一张图:一招理解LSTM/GRU门控机制
机器之心
16+阅读 · 2018年12月18日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
边缘计算的军事应用
专知会员服务
7+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
9+阅读 · 8月8日
《多域冲突比较支持模型》60页
专知会员服务
14+阅读 · 8月7日
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员