Reinforcement learning (RL) has become a dominant post-training paradigm, enabling large language models (LLMs) to learn from rewards. We observe that societal regulations are structurally similar to reward functions. They define measurable outcomes, thresholds, and exceptions, while often leaving institutional intent only partially specified. We hypothesise that the RL training process may exploit these gaps and therefore ask whether models' well-known tendency to hack reward functions during RL can scale into a more consequential failure mode named societal hacking: discovering loopholes in the rules society runs on. To study this phenomenon, we introduce SocioHack, a sandbox of 72 societal environments, and find that within these environments, reward hacking naturally emerges and leads to regulatory loophole discovery. Models learn to hack the social rules and generate strategies that remain technically compliant while defeating regulatory intent, and current LLM safeguards provide only limited mitigation. Therefore, collecting in-the-wild feedback for model training requires greater caution, and we need a next-generation post-training paradigm for safely iterating LLMs in real society.=


翻译:强化学习已成为主导性后训练范式,使大语言模型能够从奖励信号中学习。我们观察到社会规制体系与奖励函数具有结构性相似:它们既定义可量化目标、阈值和例外条款,又往往仅部分明确制度意图。我们假设强化学习训练过程可能利用这些结构缝隙,因而探究语言模型在强化学习中众所周知的奖励钻营倾向,能否升级为更具破坏性的故障模式——即"社会体系钻营":发现社会运行规则中的漏洞。为研究该现象,我们构建了包含72个社会环境的SocioHack沙盒实验平台,发现奖励钻营在这些环境中自然涌现并导致规制漏洞发掘。模型学习如何破解社会规则,生成在技术上合规却瓦解规制意图的策略,而现有语言模型防护措施仅能提供有限缓解。因此,使用自然场景反馈进行模型训练需更高审慎度,亟需发展新一代后训练范式,以实现语言模型在真实社会中的安全迭代。

0
下载
关闭预览

相关内容

大语言模型持续学习:方法、挑战与机遇
专知会员服务
22+阅读 · 3月16日
【博士论文】强化学习智能体的奖励函数设计
专知会员服务
49+阅读 · 2025年4月8日
大语言模型评估技术研究进展
专知会员服务
49+阅读 · 2024年7月9日
大型语言模型:原理、实现与发展
专知会员服务
102+阅读 · 2023年11月28日
《大语言模型进展》69页ppt,谷歌研究科学家Jason Wei
专知会员服务
88+阅读 · 2022年10月29日
「知识增强预训练语言模型」最新研究综述
专知
18+阅读 · 2022年11月18日
自然语言处理中的语言模型预训练方法
PaperWeekly
14+阅读 · 2018年10月21日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
VIP会员
最新内容
驱动军事决策变革的顶尖人工智能指挥系统
专知会员服务
5+阅读 · 8月11日
非对称防御中的自组织临界性:俄乌战争
专知会员服务
9+阅读 · 8月10日
《战争中的大语言模型监管》
专知会员服务
11+阅读 · 8月10日
《边缘计算关键技术分析及美军作战实践应用》
边缘计算的军事应用
专知会员服务
12+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
13+阅读 · 8月8日
相关基金
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
Top
微信扫码咨询专知VIP会员