Agentic reinforcement learning (RL) trains large language models to use tools, but its impact on alignment is poorly understood. We study how agentic RL for search affects the alignment of instruction-tuned (IT) models. We find that RL-trained models inherit refusal reasoning by deflecting harmful requests into benign search queries, but this breaks down under a simple diagnostic trigger that elicits a search call before refusal can occur. Under this condition, RL models produce multi-step unsafe search actions and reasoning, reducing search query safety by up to 68.6% in Qwen and Llama models relative to their IT counterparts. The effect generalises across model families, scales, and RL algorithms. To understand why, we identify linear directions in the residual stream that control search query safety, and show that RL training progressively shifts search behaviour toward the harmful end of this direction. We thus propose representation-guided RL training, which adds a reward penalty based on projection toward the harmful search direction. Training on benign data alone, it restores IT-level alignment without reducing task accuracy and requires no additional training data. Together, our work provides the first framework for diagnosing, mechanistically analysing, and mitigating alignment degradation in agentic RL for search.


翻译:智能体强化学习(Agentic RL)训练大型语言模型使用工具,但其对对齐能力的影响尚不明确。我们研究了面向搜索的智能体强化学习如何影响指令微调(IT)模型的对齐效果。研究发现,经过强化学习训练的模型会通过将有害请求转化为无害搜索查询来继承拒绝推理机制,但这一机制在诊断性触发条件下会被破坏——该条件在拒绝发生前诱发搜索调用。在此条件下,RL模型会产生多步不安全的搜索行动与推理过程,相较于对应的IT模型,Qwen和Llama系列模型的搜索查询安全性最多下降68.6%。该效应在不同模型家族、规模及强化学习算法中普遍存在。为探究成因,我们识别出残差流中控制搜索查询安全性的线性方向,并证明RL训练会逐步将搜索行为向该方向的有害端偏移。据此我们提出表征引导的强化学习训练方法,通过基于有害搜索方向投影的奖励惩罚项,仅利用良性数据即可在保持任务准确率的前提下恢复IT级对齐,且无需额外训练数据。本研究首次构建了面向搜索的智能体强化学习中对齐退化的诊断、机制分析与缓解框架。

0
下载
关闭预览

相关内容

互联网
大语言模型智能体强化学习:全景综述
专知会员服务
51+阅读 · 2025年12月18日
面向大语言模型的智能体化强化学习图景:综述
专知会员服务
56+阅读 · 2025年9月3日
面向视觉的强化学习综述
专知会员服务
23+阅读 · 2025年8月12日
《改进单智能体和多智能体深度强化学习方法》219页
专知会员服务
64+阅读 · 2025年2月14日
自动驾驶中的多智能体强化学习综述
专知会员服务
48+阅读 · 2024年8月20日
《指挥和控制强化学习智能体的对抗性攻击》
专知会员服务
72+阅读 · 2024年7月6日
「基于通信的多智能体强化学习」 进展综述
基于模型的强化学习综述
专知
42+阅读 · 2022年7月13日
【MIT博士论文】数据高效强化学习,176页pdf
综述| 当图神经网络遇上强化学习
图与推荐
35+阅读 · 2022年7月1日
探索(Exploration)还是利用(Exploitation)?强化学习如何tradeoff?
深度强化学习实验室
13+阅读 · 2020年8月23日
PlaNet 简介:用于强化学习的深度规划网络
谷歌开发者
13+阅读 · 2019年3月16日
关于强化学习(附代码,练习和解答)
深度学习
38+阅读 · 2018年1月30日
【强化学习】强化学习+深度学习=人工智能
产业智能官
55+阅读 · 2017年8月11日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2015年12月31日
国家自然科学基金
13+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
19+阅读 · 2012年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
12+阅读 · 2008年12月31日
VIP会员
最新内容
从采集到决策:美军视角下的战术情报范式重构
专知会员服务
1+阅读 · 今天2:42
《履带式无人地面战车技术发展现状》
专知会员服务
2+阅读 · 今天1:46
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
相关VIP内容
大语言模型智能体强化学习:全景综述
专知会员服务
51+阅读 · 2025年12月18日
面向大语言模型的智能体化强化学习图景:综述
专知会员服务
56+阅读 · 2025年9月3日
面向视觉的强化学习综述
专知会员服务
23+阅读 · 2025年8月12日
《改进单智能体和多智能体深度强化学习方法》219页
专知会员服务
64+阅读 · 2025年2月14日
自动驾驶中的多智能体强化学习综述
专知会员服务
48+阅读 · 2024年8月20日
《指挥和控制强化学习智能体的对抗性攻击》
专知会员服务
72+阅读 · 2024年7月6日
相关资讯
「基于通信的多智能体强化学习」 进展综述
基于模型的强化学习综述
专知
42+阅读 · 2022年7月13日
【MIT博士论文】数据高效强化学习,176页pdf
综述| 当图神经网络遇上强化学习
图与推荐
35+阅读 · 2022年7月1日
探索(Exploration)还是利用(Exploitation)?强化学习如何tradeoff?
深度强化学习实验室
13+阅读 · 2020年8月23日
PlaNet 简介:用于强化学习的深度规划网络
谷歌开发者
13+阅读 · 2019年3月16日
关于强化学习(附代码,练习和解答)
深度学习
38+阅读 · 2018年1月30日
【强化学习】强化学习+深度学习=人工智能
产业智能官
55+阅读 · 2017年8月11日
相关基金
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2015年12月31日
国家自然科学基金
13+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
19+阅读 · 2012年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
12+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员