Multi-turn agents that plan, invoke tools, and interact with environments offer a promising paradigm for solving complex tasks, yet their capabilities typically rely on very large models whose inference cost is prohibitive in practice.On-Policy Distillation (OPD) is a natural recipe for transferring such capabilities to smaller students, but we find that it suffers a characteristic failure mode in this setting: small student errors compound across turns and push the trajectory out of the teacher's familiar state distribution, so the teacher's supervision becomes least reliable precisely where the student needs it most.We propose Guided On-Policy Distillation (Guided-OPD), a simple yet effective algorithm that mixes teacher- and student-generated turns within each rollout and schedules the teacher's intervention probability along a curriculum that decays to zero.Strong guidance keeps early trajectories close to the teacher distribution and is then gradually withdrawn to recover the purely on-policy regime used at inference.On ALFWorld, ScienceWorld, and WebShop, distilling Qwen3 students from a Qwen3-30B-A3B teacher, Guided-OPD improves Score by 21.1\% and Success Rate by 25.5\% over vanilla OPD on average, with larger gains on smaller students.


翻译:多轮智能体通过规划、调用工具并与环境交互,为解决复杂任务提供了有前景的范式,但其能力通常依赖规模极大的模型,导致推理成本在实践中难以承受。在线策略蒸馏(On-Policy Distillation, OPD)是将此类能力迁移至较小学生模型的有效策略,但我们发现该方法在此场景中存在典型失效模式:学生模型在轮次间累积的小误差会将轨迹推离教师模型的熟悉状态分布,导致教师模型在最需要监督的环节提供最不可靠的指导。我们提出引导式在线策略蒸馏(Guided-OPD),一种简洁高效的算法,该算法在每个轨迹生成轮次中混合教师与学生生成的交互步骤,并按照逐步衰减至零的课程式规划策略调度教师干预概率。强引导机制使早期轨迹紧贴教师分布,随后逐步撤销引导以恢复推理阶段使用的纯在线策略模式。在ALFWorld、ScienceWorld及WebShop数据集上,通过将Qwen3-30B-A3B教师模型蒸馏至Qwen3学生模型,Guided-OPD相较于原始OPD方法,平均得分提升21.1%,成功率提升25.5%,且学生模型规模越小提升幅度越大。

0
下载
关闭预览

相关内容

综述 | OPSD:大语言模型的在线策略自蒸馏
专知会员服务
10+阅读 · 6月1日
大语言模型同策略蒸馏研究综述
专知会员服务
20+阅读 · 4月5日
《基于Transformer的智能体的战术决策解释》
专知会员服务
50+阅读 · 2025年12月28日
《分布式多智能体强化学习策略的可解释性研究》
专知会员服务
31+阅读 · 2025年11月17日
《多智能体强化学习策略优化算法设计》226页
专知会员服务
68+阅读 · 2024年6月9日
【NeurIPS 2021】设置多智能体策略梯度的方差
专知会员服务
22+阅读 · 2021年10月24日
模型压缩 | 知识蒸馏经典解读
AINLP
11+阅读 · 2020年5月31日
多智能体强化学习(MARL)近年研究概览
PaperWeekly
39+阅读 · 2020年3月15日
群体智能:新一代人工智能的重要方向
走向智能论坛
12+阅读 · 2017年8月16日
国家自然科学基金
44+阅读 · 2015年12月31日
国家自然科学基金
20+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
47+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
19+阅读 · 2008年12月31日
Arxiv
0+阅读 · 6月16日
Arxiv
0+阅读 · 6月14日
VIP会员
最新内容
致命七类无人机:无人机时代的演进型合成兵种
《异构无人水面艇集群作战自主制导算法》130页
《人工智能能通过美国陆军战争学院吗?》报告
军事域人工智能驱动系统的治理
专知会员服务
4+阅读 · 9月14日
相关VIP内容
综述 | OPSD:大语言模型的在线策略自蒸馏
专知会员服务
10+阅读 · 6月1日
大语言模型同策略蒸馏研究综述
专知会员服务
20+阅读 · 4月5日
《基于Transformer的智能体的战术决策解释》
专知会员服务
50+阅读 · 2025年12月28日
《分布式多智能体强化学习策略的可解释性研究》
专知会员服务
31+阅读 · 2025年11月17日
《多智能体强化学习策略优化算法设计》226页
专知会员服务
68+阅读 · 2024年6月9日
【NeurIPS 2021】设置多智能体策略梯度的方差
专知会员服务
22+阅读 · 2021年10月24日
相关基金
国家自然科学基金
44+阅读 · 2015年12月31日
国家自然科学基金
20+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
47+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
19+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员