Reasoning large language models exhibit complex reasoning behaviors via extended chain-of-thought generation that are highly fragile to information loss during decoding, creating critical challenges for KV cache compression. Existing token-dropping methods directly disrupt reasoning chains by removing intermediate steps, while head-reallocation methods, designed for retrieval tasks, fail to preserve the heads essential for generative reasoning. However, no existing method can identify which attention heads genuinely maintain reasoning consistency and control generation termination. To address this, we propose RLKV, which uses reinforcement learning as a probe to discover which heads contribute to reasoning quality by directly optimizing their cache usage against actual generation outcomes. This discovery naturally leads to an efficient compression strategy: we allocate full KV cache to reasoning-critical heads while aggressively compressing others with constant-size KV cache. Experiments reveal that a fraction of heads proves essential for reasoning, enabling 20--60% cache reduction with near-lossless performance across diverse tasks and models, and up to 2.06x end-to-end speedup at 60% reduction.


翻译:推理型大语言模型通过扩展的思维链生成展现复杂的推理行为,但其解码过程中对信息损失高度敏感,这为KV缓存压缩带来了关键挑战。现有的词元丢弃方法直接移除中间步骤,会破坏推理链;而面向检索任务设计的头部重分配方法,无法保留生成式推理所必需的注意力头。然而,当前尚无方法能识别哪些注意力头真正维持推理连贯性并控制生成终止。为解决这一问题,我们提出RLKV方法,利用强化学习作为探针,通过直接优化注意力头的缓存使用对实际生成结果的影响,来发现哪些头部对推理质量有贡献。这一发现自然引出了高效的压缩策略:为推理关键头部分配完整KV缓存,同时对其他头部采用固定大小的KV缓存进行激进压缩。实验表明,仅需少量头部即可保证推理效果,在多种任务和模型上实现20%-60%的缓存压缩且性能近乎无损,在60%压缩率下端到端加速比可达2.06倍。

0
下载
关闭预览

相关内容

博士论文 | 面向大模型推理的内存高效算法
专知会员服务
7+阅读 · 7月27日
大模型推理的天花板在哪里?
专知会员服务
16+阅读 · 2025年6月12日
复杂推理与慢思考
专知会员服务
49+阅读 · 2025年3月11日
TransMLA:多头潜在注意力(MLA)即为所需
专知会员服务
23+阅读 · 2025年2月13日
大模型的模型压缩与有效推理综述
专知会员服务
43+阅读 · 2024年7月8日
大型语言模型的模型压缩与高效推理:综述
专知会员服务
94+阅读 · 2024年2月17日
通过集成 XNNPACK 实现推理速度飞跃
TensorFlow
26+阅读 · 2020年7月30日
如何设计基于深度学习的图像压缩算法
论智
41+阅读 · 2018年4月26日
基础 | 基于注意力机制的seq2seq网络
黑龙江大学自然语言处理实验室
16+阅读 · 2018年3月7日
深度学习中的注意力机制
CSDN大数据
24+阅读 · 2017年11月2日
关系推理:基于表示学习和语义要素
计算机研究与发展
19+阅读 · 2017年8月22日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
Arxiv
0+阅读 · 6月12日
VIP会员
最新内容
《无人机蜂群:释放人类-蜂群编队的潜能》
专知会员服务
0+阅读 · 12分钟前
《战略战术化:一项综合性述评》
专知会员服务
0+阅读 · 16分钟前
美陆军-工业界协同推进反无人机系统技术发展
专知会员服务
1+阅读 · 38分钟前
《跨域指挥背景下的领导力发展》最新报告
专知会员服务
0+阅读 · 44分钟前
俄乌无人机战争的六大启示
专知会员服务
10+阅读 · 8月3日
《无人机空中监控:通信实验洞察》
专知会员服务
8+阅读 · 8月3日
从采集到决策:美军视角下的战术情报范式重构
相关VIP内容
博士论文 | 面向大模型推理的内存高效算法
专知会员服务
7+阅读 · 7月27日
大模型推理的天花板在哪里?
专知会员服务
16+阅读 · 2025年6月12日
复杂推理与慢思考
专知会员服务
49+阅读 · 2025年3月11日
TransMLA:多头潜在注意力(MLA)即为所需
专知会员服务
23+阅读 · 2025年2月13日
大模型的模型压缩与有效推理综述
专知会员服务
43+阅读 · 2024年7月8日
大型语言模型的模型压缩与高效推理:综述
专知会员服务
94+阅读 · 2024年2月17日
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员