Top-K teacher logits make on-policy distillation tractable, but probability mass is not the same as decision support. In a two-teacher tool-use setting, vanilla generalized knowledge distillation raises tool-call recall while also calling on examples that require direct answers. The response teacher's top-32 retains 99.99% of its probability mass yet contains the tool-call behavior-switch token on only 0.4% of 1,500 audited prompts; even top-256 covers only 52.2%. Because omitted logits receive zero direct gradient under the truncated objective, the tool teacher reinforces entry while the response teacher usually cannot oppose it. Frozen replay shows that a wrong entry then amplifies divergence along the generated trajectory. Restoring the tool-call token only at the first response position moves first-token entry but mostly delays eventual calls. Restoring it at every response position closes this gap: across three matched seeds, full-generation over-calling falls from 14.2+/-2.1% to 3.7+/-0.5%. The correction is not class-selective: tool-call recall falls from 91.5+/-1.7% to 79.1+/-1.7%, and exact multi-turn success from 7.0+/-0.3% to 3.4+/-0.5%. The matched interventions therefore identify decision-critical support omission as a causal mechanism of the observed drift under the tested teacher-top-32 objective, but restoring it traces a conservatism-capability trade-off rather than a free improvement. We further compare fixed clipping, global reweighting, localized soft compression, and validation-tuned inference-time bias as alternative ways to move this operating point. These results expose a failure mode hidden by near-complete probability-mass coverage and motivate support-aware auditing of compressed distillation.


翻译:暂无翻译

0
下载
关闭预览

相关内容

综述 | 多模态大模型的不确定性感知决策
专知会员服务
8+阅读 · 8月19日
【2023新书】决策支持系统和自动谈判, 240页pdf
专知会员服务
48+阅读 · 2023年6月24日
《多目标强化学习和规划的实用指南》59页最新论文
专知会员服务
56+阅读 · 2022年8月10日
【斯坦福2021新书】决策算法,694页pdf阐述不确定性决策
专知会员服务
264+阅读 · 2021年1月27日
一文理解Ranking Loss/Margin Loss/Triplet Loss
极市平台
16+阅读 · 2020年8月10日
近期语音类前沿论文
深度学习每日摘要
14+阅读 · 2019年3月17日
半监督多任务学习:Semisupervised Multitask Learning
我爱读PAMI
18+阅读 · 2018年4月29日
Focal Loss for Dense Object Detection
统计学习与视觉计算组
12+阅读 · 2018年3月15日
一文读懂「Attention is All You Need」| 附代码实现
PaperWeekly
37+阅读 · 2018年1月10日
国家自然科学基金
122+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
Arxiv
0+阅读 · 8月3日
VIP会员
相关主题
最新内容
受限仓库多智能体取送中的动态安全等待点选择
《国防技术管理》印度智库报告最新45页
专知会员服务
3+阅读 · 8月28日
《美陆军最新条令:保障行动》
专知会员服务
4+阅读 · 8月28日
算法战场:人工智能如何重新定义军事力量
专知会员服务
6+阅读 · 8月28日
《北约联邦式电子战云架构》
专知会员服务
6+阅读 · 8月27日
《美陆军野战手册:空域管理战术》
专知会员服务
10+阅读 · 8月27日
相关基金
国家自然科学基金
122+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员