We introduce Gram, an automated alignment auditing framework to assess the propensity of AI agents to engage in sabotage. We evaluate Gemini models across 17 simulated agentic deployment scenarios that incentivize sabotage. We find Gemini models misbehave in about 2-3% of our simulated trajectories. Many of these cases are explained by "overeagerness" in Gemini models resulting in both excessive role-playing and goal-seeking behavior. In contrast to other alignment auditing approaches, Gram is designed to specifically evaluate misalignment and intentional sabotage in agentic coding and research agents. We additionally introduce an experimental investigator agent pipeline which enables fine-grained targeted experiments to identify the drivers of misbehavior. We find that increasing realism of environments and removing nudges to misbehave tends to reduce sabotage rates close to zero.
翻译:我们介绍了Gram,一个用于评估AI代理参与破坏倾向的自动化对齐审计框架。我们在17个模拟代理部署场景下评估了Gemini模型,这些场景激励了破坏行为。我们发现Gemini模型在我们的模拟轨迹中约有2-3%出现不当行为。其中许多案例可归因于Gemini模型中的“过度热切”现象,导致过度角色扮演和目标追求行为。与其他对齐审计方法相比,Gram专门设计用于评估代理型编码和研究代理中的失调与蓄意破坏。我们还引入了一个实验性调查代理流水线,能够进行细粒度定向实验以识别不当行为的驱动因素。我们发现,提高环境真实性并移除不当行为的诱导因素可将破坏率降至接近零。