Large language models are increasingly deployed as coding agents, shifting safety from individual responses to action sequences. Existing benchmarks, however, primarily assess whether models refuse unsafe prompts, leaving impacts on stateful workspaces largely unexamined. We present SABER, a benchmark for environment-aware operational safety that places models in realistic agent-style projects and evaluates safety from the final environment state after a sequence of actions. Beyond binary safety-violation reports, SABER categorizes violations by cause, enabling analysis of model-specific safety profiles. Our evaluations show that even the best-performing model has more than a 54% harmful safety-violation rate (HSR), suggesting that current alignment remains insufficient for realistic project environments. SABER further reveals distinct safety profiles across models. Our benchmark is publicly available at https://github.com/sssr-lab/saber.
翻译:大型语言模型正越来越多地被部署为编码体,其安全性从个体响应转向了动作序列。然而,现有基准主要评估模型是否拒绝不安全提示,而忽略了动作对有状态工作空间的影响。我们提出SABER,一个面向环境感知操作安全的基准,该基准将模型置于真实的体风格项目中,并根据一系列动作后的最终环境状态评估安全性。除了二元的安全违规报告,SABER还按原因对违规进行分类,从而能够分析模型特定的安全概况。我们的评估显示,即使性能最佳的模型,其有害安全违规率(HSR)也超过54%,这表明当前的对齐机制对于真实项目环境仍显不足。SABER还揭示了不同模型间差异化的安全概况。我们的基准已在 https://github.com/sssr-lab/saber 公开提供。