UK AISI Alignment Evaluation Case-Study

This technical report presents methods developed by the UK AI Security Institute for assessing whether advanced AI systems reliably follow intended goals. Specifically, we evaluate whether frontier models sabotage safety research when deployed as coding assistants within an AI lab. Applying our methods to four frontier models, we find no confirmed instances of research sabotage. However, we observe that Claude Opus 4.5 Preview (a pre-release snapshot of Opus 4.5) and Sonnet 4.5 frequently refuse to engage with safety-relevant research tasks, citing concerns about research direction, involvement in self-training, and research scope. We additionally find that Opus 4.5 Preview shows reduced unprompted evaluation awareness compared to Sonnet 4.5, while both models can distinguish evaluation from deployment scenarios when prompted. Our evaluation framework builds on Petri, an open-source LLM auditing tool, with a custom scaffold designed to simulate realistic internal deployment of a coding agent. We validate that this scaffold produces trajectories that all tested models fail to reliably distinguish from real deployment data. We test models across scenarios varying in research motivation, activity type, replacement threat, and model autonomy. Finally, we discuss limitations including scenario coverage and evaluation awareness.

翻译：本技术报告介绍了英国人工智能安全研究所开发的用于评估高级人工智能系统是否可靠遵循预期目标的方法。具体而言，我们评估了前沿模型在被部署为人工智能实验室内的编程助手时，是否会破坏安全研究。将我们的方法应用于四个前沿模型，我们并未发现确凿的研究破坏案例。然而，我们观察到Claude Opus 4.5 Preview（Opus 4.5的预发布快照）和Sonnet 4.5频繁拒绝参与安全相关的研究任务，其理由涉及研究方向、参与自身训练以及研究范围。此外，我们还发现，与Sonnet 4.5相比，Opus 4.5 Preview在无提示条件下的评估意识较低，而两个模型在被提示时均能区分评估场景与部署场景。我们的评估框架基于开源的大语言模型审计工具Petri构建，并采用定制的支架来模拟编程智能体在内部部署的真实场景。我们验证了该支架生成的轨迹，所有受测模型均无法将其与真实部署数据可靠区分。我们在研究动机、活动类型、替换威胁以及模型自主性等不同场景下对模型进行了测试。最后，我们讨论了包括场景覆盖率和评估意识在内的局限性。

相关内容

MoDELS

关注 45

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

《运用人工神经网络的防空系统威胁评估模型》

专知会员服务

16+阅读 · 2月21日

前沿人工智能趋势报告（Frontier AI Trends Report）

专知会员服务

40+阅读 · 2025年12月20日

《防务领域人工智能可信赖性：为防务开发负责任、符合伦理且可信赖的AI系统》欧洲防务局2025最新107页

专知会员服务

23+阅读 · 2025年5月14日

覆盖800+文献、多位知名学者挂帅，北大联合剑桥、CMU等多所高校发布《AI 对齐 (Alignment)》全面性综述

专知会员服务

54+阅读 · 2023年11月1日