Large language models are increasingly proposed as autonomous agents for high-stakes public workflows, yet we lack systematic evidence about whether they would follow institutional rules when granted authority. We present evidence that integrity in institutional AI should be treated as a pre-deployment requirement rather than a post-deployment assumption. We evaluate multi-agent governance simulations in which agents occupy formal governmental roles under different authority structures, and we score rule-breaking and abuse outcomes with an independent rubric-based judge across 28,112 transcript segments. While we advance this position, the core contribution is empirical: among models operating below saturation, governance structure is a stronger driver of corruption-related outcomes than model identity, with large differences across regimes and model--governance pairings. Lightweight safeguards can reduce risk in some settings but do not consistently prevent severe failures. These results imply that institutional design is a precondition for safe delegation: before real authority is assigned to LLM agents, systems should undergo stress testing under governance-like constraints with enforceable rules, auditable logs, and human oversight on high-impact actions.
翻译:大型语言模型日益被提议作为高风险公共工作流程中的自主智能体,但我们缺乏系统性的证据来证明它们在获得授权后是否会遵守制度规则。我们提出证据表明,机构性AI中的诚信应被视为部署前的必要条件而非部署后的假设。我们评估了多智能体治理模拟,其中智能体在不同权力结构下担任正式政府角色,并通过一个独立的基于评分标准的评判者,对28,112个对话片段中的违规行为和权力滥用结果进行评分。在推进这一立场的同时,核心贡献是实证性的:在低于饱和状态运行的模型中,治理结构比模型身份更能驱动腐败相关结果,不同制度及模型-治理配对之间存在显著差异。轻量级的安全措施可以在某些环境中降低风险,但无法始终如一地防止严重失败。这些结果表明,制度设计是安全授权的先决条件:在将实际权力分配给LLM智能体之前,系统应在类似治理约束下进行压力测试,并配备可强制执行的规则、可审计的日志以及对高影响行动的人工监督。