Adversarial robustness evaluations of large language models (LLMs) typically report attack success rate (ASR) under fixed query budgets, implicitly treating all attacks as equally costly. In practice, the computational expense of different attack strategies can vary by orders of magnitude. Consequently, ASR at a fixed budget can obscure the true effort required to jailbreak a model, thereby making it hard to determine whether an attack's cost justifies its payoff to the attacker. We propose a compute-aware evaluation framework based on computational pressure, measured in cumulative floating-point operations (FLOPs), as a proxy for adversarial effort. We introduce risk-compute curves, which map compute budgets to attack risk, and derive two metrics that summarize the average pressure required for a given attack to succeed. Across ten models spanning three families and four different stages in language model training and alignment, evaluated with three attack strategies (gradient-based, iterative refinement, and template-based) on two jailbreak robustness benchmarks, we find: (1) alignment training has non-monotonic effects on compute-space robustness; (2) scaling model size reduces gradient-based attack effectiveness but has limited impact on cheaper template-based attacks; (3) gradient-based attacks optimized on a surrogate model can transfer to a separate target model, providing a way to reduce attacker costs; (4) compute cost varies by up to ${\approx}5{\times}$ across harm categories within a single model; and (5) safety-aligned RL increases aggregate cost while leaving some categories disproportionately accessible. We release our framework to enable compute-aware risk assessment and evaluation.
翻译:大型语言模型(LLM)的对抗鲁棒性评估通常报告在固定查询预算下的攻击成功率(ASR),隐含地视所有攻击为同等代价。实际上,不同攻击策略的计算开销可能相差数个数量级。因此,固定预算下的ASR可能掩盖破解模型所需的真实努力,从而难以确定攻击成本是否值得攻击者投入。我们提出一种基于计算压力的计算感知评估框架,以累积浮点运算次数(FLOPs)作为对抗努力的代理指标。我们引入风险-计算曲线,将计算预算映射为攻击风险,并推导出两个指标来总结给定攻击成功所需的平均压力。在横跨三个模型家族、四个不同语言模型训练与对齐阶段的十个模型上,使用三种攻击策略(梯度基、迭代细化、模板基)对两个破解鲁棒性基准进行评估,我们发现:(1)对齐训练对计算空间鲁棒性具有非单调影响;(2)扩大模型规模会降低梯度基攻击的有效性,但对更廉价的模板基攻击影响有限;(3)在代理模型上优化的梯度基攻击可以迁移至独立的目标模型,从而降低攻击者成本;(4)单个模型内不同危害类别的计算成本差异高达约5倍;(5)安全对齐的强化学习在增加总体成本的同时,使某些类别仍保持不成比例的可及性。我们发布该框架以实现计算感知的风险评估与评估。