Multi-turn agents that plan, invoke tools, and interact with environments offer a promising paradigm for solving complex tasks, yet their capabilities typically rely on very large models whose inference cost is prohibitive in practice.On-Policy Distillation (OPD) is a natural recipe for transferring such capabilities to smaller students, but we find that it suffers a characteristic failure mode in this setting: small student errors compound across turns and push the trajectory out of the teacher's familiar state distribution, so the teacher's supervision becomes least reliable precisely where the student needs it most.We propose Guided On-Policy Distillation (Guided-OPD), a simple yet effective algorithm that mixes teacher- and student-generated turns within each rollout and schedules the teacher's intervention probability along a curriculum that decays to zero.Strong guidance keeps early trajectories close to the teacher distribution and is then gradually withdrawn to recover the purely on-policy regime used at inference.On ALFWorld, ScienceWorld, and WebShop, distilling Qwen3 students from a Qwen3-30B-A3B teacher, Guided-OPD improves Score by 21.1\% and Success Rate by 25.5\% over vanilla OPD on average, with larger gains on smaller students.


翻译:多轮智能体通过规划、调用工具并与环境交互,为解决复杂任务提供了有前景的范式,但其能力通常依赖规模极大的模型,导致推理成本在实践中难以承受。在线策略蒸馏(On-Policy Distillation, OPD)是将此类能力迁移至较小学生模型的有效策略,但我们发现该方法在此场景中存在典型失效模式:学生模型在轮次间累积的小误差会将轨迹推离教师模型的熟悉状态分布,导致教师模型在最需要监督的环节提供最不可靠的指导。我们提出引导式在线策略蒸馏(Guided-OPD),一种简洁高效的算法,该算法在每个轨迹生成轮次中混合教师与学生生成的交互步骤,并按照逐步衰减至零的课程式规划策略调度教师干预概率。强引导机制使早期轨迹紧贴教师分布,随后逐步撤销引导以恢复推理阶段使用的纯在线策略模式。在ALFWorld、ScienceWorld及WebShop数据集上,通过将Qwen3-30B-A3B教师模型蒸馏至Qwen3学生模型,Guided-OPD相较于原始OPD方法,平均得分提升21.1%,成功率提升25.5%,且学生模型规模越小提升幅度越大。

0
下载
关闭预览

相关内容

综述 | OPSD:大语言模型的在线策略自蒸馏
专知会员服务
10+阅读 · 6月1日
大语言模型同策略蒸馏研究综述
专知会员服务
20+阅读 · 4月5日
《基于Transformer的智能体的战术决策解释》
专知会员服务
49+阅读 · 2025年12月28日
《分布式多智能体强化学习策略的可解释性研究》
专知会员服务
30+阅读 · 2025年11月17日
《多智能体强化学习策略优化算法设计》226页
专知会员服务
67+阅读 · 2024年6月9日
【NeurIPS 2021】设置多智能体策略梯度的方差
专知会员服务
21+阅读 · 2021年10月24日
模型压缩 | 知识蒸馏经典解读
AINLP
11+阅读 · 2020年5月31日
多智能体强化学习(MARL)近年研究概览
PaperWeekly
38+阅读 · 2020年3月15日
群体智能:新一代人工智能的重要方向
走向智能论坛
12+阅读 · 2017年8月16日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
20+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
47+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
Arxiv
0+阅读 · 6月16日
Arxiv
0+阅读 · 6月14日
VIP会员
最新内容
从采集到决策:美军视角下的战术情报范式重构
专知会员服务
0+阅读 · 今天2:42
《履带式无人地面战车技术发展现状》
专知会员服务
2+阅读 · 今天1:46
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
相关VIP内容
综述 | OPSD:大语言模型的在线策略自蒸馏
专知会员服务
10+阅读 · 6月1日
大语言模型同策略蒸馏研究综述
专知会员服务
20+阅读 · 4月5日
《基于Transformer的智能体的战术决策解释》
专知会员服务
49+阅读 · 2025年12月28日
《分布式多智能体强化学习策略的可解释性研究》
专知会员服务
30+阅读 · 2025年11月17日
《多智能体强化学习策略优化算法设计》226页
专知会员服务
67+阅读 · 2024年6月9日
【NeurIPS 2021】设置多智能体策略梯度的方差
专知会员服务
21+阅读 · 2021年10月24日
相关基金
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
20+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
47+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员