Ensuring that AI agents behave safely and beneficially when interacting with other parties has emerged as one of the central challenges of modern AI safety. While mechanism design, as the theory of designing rules to align individual and collective objectives, can incentivize cooperative behavior, it is still an open question whether it alone is sufficient to maximize LLM agents' social welfare. This work proves that the answer is negative: drawing from incomplete contract theory, we formally show that when contracts cannot distinguish all relevant future contingencies, there is a strictly positive welfare loss that no realistic mechanism can eliminate. We show that prosocial agents, who weigh others' welfare alongside their own, can close this gap and achieve outcomes that are socially superior and individually beneficial. Experimentally, we show that in multi-agent resource-allocation environments and canonical social dilemmas where agents are powered by large language models, prosociality is beneficial. The implication for AI safety is clear: to enable cooperative interactions at scale, designing adequate mechanisms is not sufficient; agents must be built to be intrinsically prosocial.


翻译:确保人工智能智能体在与其它主体交互时表现出安全且有益的行为,已成为现代人工智能安全领域的核心挑战之一。尽管机制设计作为设计规则以协调个体与集体目标的理论,能够激励合作行为,但其本身是否足以最大化大型语言模型智能体的社会总福利仍是一个开放性问题。本研究证明答案是否定的:基于不完全合同理论,我们形式化地证明,当合同无法区分所有相关的未来偶发事件时,存在任何现实机制都无法消除的正福利损失。我们表明,能够权衡他人福祉与自身利益的亲社会智能体可以弥合这一差距,实现社会更优且个体有益的结果。实验方面,我们证明在大语言模型驱动的多智能体资源分配环境及经典社会困境中,亲社会性具有正向作用。这对人工智能安全的启示是明确的:要实现大规模的合作交互,设计充分的机制并不足够;智能体必须被构建为具备内在亲社会性。

0
下载
关闭预览

相关内容

设计是对现有状的一种重新认识和打破重组的过程,设计让一切变得更美。
AI智能体基础设施
专知会员服务
44+阅读 · 2025年7月12日
《多智能体强化学习中的机制设计优化研究》103页
专知会员服务
34+阅读 · 2025年5月31日
《在单智能体与多智能体AI系统中融入人类合理性》100页
《多智能体强化学习中机制设计的优化》103页
专知会员服务
32+阅读 · 2025年5月3日
AI智能体面临的威胁:关键安全挑战与未来路径综述
专知会员服务
53+阅读 · 2024年6月7日
《人工智能辅助决策面临的三大挑战》
专知会员服务
87+阅读 · 2023年12月15日
《结合机器人行为以实现安全、智能的执行》
专知会员服务
17+阅读 · 2023年7月4日
【人机融合智能】人机融合智能的现状与展望
产业智能官
12+阅读 · 2020年3月18日
浅谈群体智能——新一代AI的重要方向
中国科学院自动化研究所
44+阅读 · 2019年10月16日
面向人工智能的计算机体系结构
计算机研究与发展
14+阅读 · 2019年6月6日
【混合智能】人机混合智能的哲学思考
产业智能官
12+阅读 · 2018年10月28日
群体智能:新一代人工智能的重要方向
走向智能论坛
12+阅读 · 2017年8月16日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
Arxiv
0+阅读 · 5月28日
VIP会员
最新内容
俄乌无人机战争的六大启示
专知会员服务
4+阅读 · 今天7:07
《无人机空中监控:通信实验洞察》
专知会员服务
3+阅读 · 今天7:05
从采集到决策:美军视角下的战术情报范式重构
《履带式无人地面战车技术发展现状》
专知会员服务
6+阅读 · 8月2日
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
6+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
13+阅读 · 7月31日
相关VIP内容
AI智能体基础设施
专知会员服务
44+阅读 · 2025年7月12日
《多智能体强化学习中的机制设计优化研究》103页
专知会员服务
34+阅读 · 2025年5月31日
《在单智能体与多智能体AI系统中融入人类合理性》100页
《多智能体强化学习中机制设计的优化》103页
专知会员服务
32+阅读 · 2025年5月3日
AI智能体面临的威胁:关键安全挑战与未来路径综述
专知会员服务
53+阅读 · 2024年6月7日
《人工智能辅助决策面临的三大挑战》
专知会员服务
87+阅读 · 2023年12月15日
《结合机器人行为以实现安全、智能的执行》
专知会员服务
17+阅读 · 2023年7月4日
相关基金
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
Top
微信扫码咨询专知VIP会员