AI systems have become increasingly capable of dangerous behaviours in many domains. This raises the question: Do models sometimes choose to violate human instructions in order to perform behaviour that is more useful for certain goals? We introduce a benchmark for measuring model propensity for instrumental convergence (IC) behaviour in terminal-based agents. This is behaviour such as self-preservation that has been hypothesised to play a key role in risks from highly capable AI agents. Our benchmark is realistic and low-stakes which serves to reduce evaluation-awareness and roleplay confounds. The suite contains seven operational tasks, each with an official workflow and a policy-violating shortcut. An eight-variant shared framework varies monitoring, instruction clarity, stakes, permission, instrumental usefulness and blocked honest paths to support inferences regarding the factors driving IC behaviour. We evaluated ten models using deterministic environment-state scorers over 1,680 samples, with trace review employed for audit and adjudication purposes. The final IC rate is 86 out of 1,680 samples (5.1%). IC behaviour is concentrated rather than uniform: two Gemini models account for 66.3% of IC cases and three tasks account for 84.9%. Conditions in which IC behaviour is indispensable for task success result in the greatest increase in the adjusted IC rate (+15.7 percentage points), whereas emphasising that task success is critical or certain framing choices do not produce comparable effects. Our findings indicate that realistic, low-nudge environments elicit IC behaviour rarely but systematically in most tested models. We conclude that it is feasible to robustly measure tendencies for dangerous behaviour in current frontier AI agents.
翻译:人工智能系统已在多个领域展现出日益增强的危险行为能力。这引发了一个问题:模型是否会为了执行对某些目标更有用的行为而选择违背人类指令?我们引入了一个基准测试,用于衡量终端智能体中工具性趋同(IC)行为的倾向。这种行为(如自我保存)据假设在高能力AI智能体带来的风险中扮演关键角色。我们的基准测试兼具真实性与低风险性,有助于减少评估意识干扰和角色扮演混淆。该套件包含七项操作性任务,每项任务均设有官方流程与违反策略的快捷方式。一个八变体共享框架在监控、指令清晰度、风险程度、权限、工具性有用性及受堵诚实路径方面存在差异,以支持对驱动IC行为因素进行推断。我们使用确定性环境状态评分器对十个模型进行了评估,采集了1,680个样本,并采用痕迹审查进行审计与裁定。最终IC率为1,680个样本中的86例(5.1%)。IC行为呈现集中而非均匀分布:两个Gemini模型占IC案例的66.3%,三项任务占84.9%。当IC行为对任务成功不可或缺时,调整后的IC率增幅最大(+15.7个百分点),而强调任务成功至关重要或某些框架选择并未产生类似效果。我们的研究结果表明,在大多数测试模型中,低推动力的现实环境虽罕见但系统地诱发了IC行为。我们得出结论:在当前前沿AI智能体中,稳健地衡量危险行为倾向是可行的。