AI-control monitors score individual agent actions to detect misbehavior, but real harm can be distributed across many benign-looking steps, each individually below any per-step alarm. We construct a marginal-preserving, correlation-encoded distributed-sabotage attack using a Gaussian-copula AR(1) construction: the per-step monitor-score marginal is held exactly equal to benign, so mean, max, top-k tail, and threshold monitors (Monitor A) are defeated by construction, while harm is encoded in the temporal correlation structure. We sequence the paper around three reviewer-mandated gates. (1) Realizability gate: the stealthy attack achieves KS-distance to benign of 0.013 (effectively zero) at all tested harm levels up to 3.0, confirming that harm is fully decoupled from the per-step marginal and realizability is not harm-limited. (2) Monitor-A-vs-B reconciliation: we show formally that the attack, built against Monitor A's score marginal, remains marginal-preserving under a different-score Monitor B (the correlation/sequence family: CUSUM, SPRT, HMM-LR, runs test, autocorrelation, windowed logistic), and scope worst-case claims to score functions that admit a temporal signature. (3) Non-empty detectability band: Monitor A achieves AUC 0.52 (chance); Monitor B spans AUC 0.79-0.97 at the same 1% FPR target, and as harm is amortized over more steps Monitor A collapses to chance while Monitor B holds at AUC ~0.95. These results demonstrate a non-empty detectability band and characterize the sub-threshold sabotage frontier: distribution-shape monitors fail by construction; temporal-correlation monitors can detect but are not trivially optimal.
翻译:AI监控系统通过评估个体智能体动作来检测异常行为,但实际危害可能分散在多个看似正常的步骤中,每个步骤均低于单步报警阈值。我们构建了一种基于高斯连接函数AR(1)结构的边际保持、相关性编码分布式破坏攻击:单步监控得分边际与正常行为完全一致,因此均值、最大值、top-k尾部和阈值监控器(A类监控器)在设计上失效,而危害被编码在时间相关性结构中。我们围绕三个审稿人设定的关卡组织本文:(1)可实现性关卡:在高达3.0的所有测试危害水平下,隐蔽攻击与正常行为的KS距离仅为0.013(实际为零),证实危害与单步边际完全解耦,且可实现性不受危害限制;(2)A类与B类监控器协调:我们形式上证明,针对A类监控器得分边际构建的攻击在采用不同得分的B类监控器(相关性/序列类:CUSUM、SPRT、HMM-LR、游程检验、自相关、窗口逻辑)下仍保持边际不变,并将最坏情况声明限定于存在时间特征痕迹的得分函数;(3)非空可检测带:A类监控器AUC为0.52(随机水平);在相同1%假阳性率目标下,B类监控器AUC范围为0.79-0.97,且当危害被分摊至更多步骤时,A类监控器降至随机水平,而B类监控器保持AUC约0.95。这些结果证明存在非空可检测带,并刻画了亚阈值破坏前沿:分布形态监控器因设计限制而失效;时间相关性监控器虽能检测但并非天然最优。