Agentic reinforcement learning (RL) trains large language models to use tools, but its impact on alignment is poorly understood. We study how agentic RL for search affects the alignment of instruction-tuned (IT) models. We find that RL-trained models inherit refusal reasoning by deflecting harmful requests into benign search queries, but this breaks down under a simple diagnostic trigger that elicits a search call before refusal can occur. Under this condition, RL models produce multi-step unsafe search actions and reasoning, reducing search query safety by up to 68.6% in Qwen and Llama models relative to their IT counterparts. The effect generalises across model families, scales, and RL algorithms. To understand why, we identify linear directions in the residual stream that control search query safety, and show that RL training progressively shifts search behaviour toward the harmful end of this direction. We thus propose representation-guided RL training, which adds a reward penalty based on projection toward the harmful search direction. Training on benign data alone, it restores IT-level alignment without reducing task accuracy and requires no additional training data. Together, our work provides the first framework for diagnosing, mechanistically analysing, and mitigating alignment degradation in agentic RL for search.


翻译:智能体强化学习(Agentic RL)训练大型语言模型使用工具,但其对对齐能力的影响尚不明确。我们研究了面向搜索的智能体强化学习如何影响指令微调(IT)模型的对齐效果。研究发现,经过强化学习训练的模型会通过将有害请求转化为无害搜索查询来继承拒绝推理机制,但这一机制在诊断性触发条件下会被破坏——该条件在拒绝发生前诱发搜索调用。在此条件下,RL模型会产生多步不安全的搜索行动与推理过程,相较于对应的IT模型,Qwen和Llama系列模型的搜索查询安全性最多下降68.6%。该效应在不同模型家族、规模及强化学习算法中普遍存在。为探究成因,我们识别出残差流中控制搜索查询安全性的线性方向,并证明RL训练会逐步将搜索行为向该方向的有害端偏移。据此我们提出表征引导的强化学习训练方法,通过基于有害搜索方向投影的奖励惩罚项,仅利用良性数据即可在保持任务准确率的前提下恢复IT级对齐,且无需额外训练数据。本研究首次构建了面向搜索的智能体强化学习中对齐退化的诊断、机制分析与缓解框架。

0
下载
关闭预览

相关内容

互联网
大语言模型智能体强化学习:全景综述
专知会员服务
52+阅读 · 2025年12月18日
面向大语言模型的智能体化强化学习图景:综述
专知会员服务
57+阅读 · 2025年9月3日
面向视觉的强化学习综述
专知会员服务
23+阅读 · 2025年8月12日
《改进单智能体和多智能体深度强化学习方法》219页
专知会员服务
64+阅读 · 2025年2月14日
自动驾驶中的多智能体强化学习综述
专知会员服务
49+阅读 · 2024年8月20日
《指挥和控制强化学习智能体的对抗性攻击》
专知会员服务
73+阅读 · 2024年7月6日
「基于通信的多智能体强化学习」 进展综述
基于模型的强化学习综述
专知
42+阅读 · 2022年7月13日
【MIT博士论文】数据高效强化学习,176页pdf
综述| 当图神经网络遇上强化学习
图与推荐
35+阅读 · 2022年7月1日
探索(Exploration)还是利用(Exploitation)?强化学习如何tradeoff?
深度强化学习实验室
13+阅读 · 2020年8月23日
PlaNet 简介:用于强化学习的深度规划网络
谷歌开发者
13+阅读 · 2019年3月16日
关于强化学习(附代码,练习和解答)
深度学习
38+阅读 · 2018年1月30日
【强化学习】强化学习+深度学习=人工智能
产业智能官
55+阅读 · 2017年8月11日
国家自然科学基金
44+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2015年12月31日
国家自然科学基金
13+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
19+阅读 · 2012年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
12+阅读 · 2008年12月31日
VIP会员
最新内容
《异构无人水面艇集群作战自主制导算法》130页
《人工智能能通过美国陆军战争学院吗?》报告
军事域人工智能驱动系统的治理
专知会员服务
4+阅读 · 9月14日
相关VIP内容
大语言模型智能体强化学习:全景综述
专知会员服务
52+阅读 · 2025年12月18日
面向大语言模型的智能体化强化学习图景:综述
专知会员服务
57+阅读 · 2025年9月3日
面向视觉的强化学习综述
专知会员服务
23+阅读 · 2025年8月12日
《改进单智能体和多智能体深度强化学习方法》219页
专知会员服务
64+阅读 · 2025年2月14日
自动驾驶中的多智能体强化学习综述
专知会员服务
49+阅读 · 2024年8月20日
《指挥和控制强化学习智能体的对抗性攻击》
专知会员服务
73+阅读 · 2024年7月6日
相关资讯
「基于通信的多智能体强化学习」 进展综述
基于模型的强化学习综述
专知
42+阅读 · 2022年7月13日
【MIT博士论文】数据高效强化学习,176页pdf
综述| 当图神经网络遇上强化学习
图与推荐
35+阅读 · 2022年7月1日
探索(Exploration)还是利用(Exploitation)?强化学习如何tradeoff?
深度强化学习实验室
13+阅读 · 2020年8月23日
PlaNet 简介:用于强化学习的深度规划网络
谷歌开发者
13+阅读 · 2019年3月16日
关于强化学习(附代码,练习和解答)
深度学习
38+阅读 · 2018年1月30日
【强化学习】强化学习+深度学习=人工智能
产业智能官
55+阅读 · 2017年8月11日
相关基金
国家自然科学基金
44+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2015年12月31日
国家自然科学基金
13+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
19+阅读 · 2012年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
12+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员