Large language models are increasingly proposed as autonomous agents for high-stakes public workflows, yet we lack systematic evidence about whether they would follow institutional rules when granted authority. We present evidence that integrity in institutional AI should be treated as a pre-deployment requirement rather than a post-deployment assumption. We evaluate multi-agent governance simulations in which agents occupy formal governmental roles under different authority structures, and we score rule-breaking and abuse outcomes with an independent rubric-based judge across 28,112 transcript segments. While we advance this position, the core contribution is empirical: among models operating below saturation, governance structure is a stronger driver of corruption-related outcomes than model identity, with large differences across regimes and model--governance pairings. Lightweight safeguards can reduce risk in some settings but do not consistently prevent severe failures. These results imply that institutional design is a precondition for safe delegation: before real authority is assigned to LLM agents, systems should undergo stress testing under governance-like constraints with enforceable rules, auditable logs, and human oversight on high-impact actions.


翻译:大型语言模型日益被提议作为高风险公共工作流程中的自主智能体,但我们缺乏系统性的证据来证明它们在获得授权后是否会遵守制度规则。我们提出证据表明,机构性AI中的诚信应被视为部署前的必要条件而非部署后的假设。我们评估了多智能体治理模拟,其中智能体在不同权力结构下担任正式政府角色,并通过一个独立的基于评分标准的评判者,对28,112个对话片段中的违规行为和权力滥用结果进行评分。在推进这一立场的同时,核心贡献是实证性的:在低于饱和状态运行的模型中,治理结构比模型身份更能驱动腐败相关结果,不同制度及模型-治理配对之间存在显著差异。轻量级的安全措施可以在某些环境中降低风险,但无法始终如一地防止严重失败。这些结果表明,制度设计是安全授权的先决条件:在将实际权力分配给LLM智能体之前,系统应在类似治理约束下进行压力测试,并配备可强制执行的规则、可审计的日志以及对高影响行动的人工监督。

0
下载
关闭预览

相关内容

《多智能体大语言模型系统的可靠决策研究》
专知会员服务
41+阅读 · 2月2日
智能体评判者(Agent-as-a-Judge)研究综述
专知会员服务
37+阅读 · 1月9日
基于大模型的智能体中由自主性引发的安全风险综述
专知会员服务
18+阅读 · 2025年7月1日
面向大模型多智能体系统的多维评估方法
专知会员服务
35+阅读 · 2025年4月15日
先进人工智能的多智能体风险
专知会员服务
28+阅读 · 2025年2月22日
《人工智能安全测评白皮书》,99页pdf
专知
36+阅读 · 2022年2月26日
智能时代如何构建金融反欺诈体系?
数据猿
12+阅读 · 2018年3月26日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
Arxiv
24+阅读 · 2024年2月23日
VIP会员
最新内容
对抗环境下超视距目标打击的情报支援
专知会员服务
3+阅读 · 今天14:49
《无人机对海面作战影响评估》
专知会员服务
11+阅读 · 7月21日
印度精确打击与指挥架构的断层
专知会员服务
6+阅读 · 7月20日
美空军AI完成F-16战斗机自主空战历史性试飞
专知会员服务
6+阅读 · 7月20日
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
Top
微信扫码咨询专知VIP会员