AI-control monitors score individual agent actions to detect misbehavior, but real harm can be distributed across many benign-looking steps, each individually below any per-step alarm. We construct a marginal-preserving, correlation-encoded distributed-sabotage attack using a Gaussian-copula AR(1) construction: the per-step monitor-score marginal is held exactly equal to benign, so mean, max, top-k tail, and threshold monitors (Monitor A) are defeated by construction, while harm is encoded in the temporal correlation structure. We sequence the paper around three reviewer-mandated gates. (1) Realizability gate: the stealthy attack achieves KS-distance to benign of 0.013 (effectively zero) at all tested harm levels up to 3.0, confirming that harm is fully decoupled from the per-step marginal and realizability is not harm-limited. (2) Monitor-A-vs-B reconciliation: we show formally that the attack, built against Monitor A's score marginal, remains marginal-preserving under a different-score Monitor B (the correlation/sequence family: CUSUM, SPRT, HMM-LR, runs test, autocorrelation, windowed logistic), and scope worst-case claims to score functions that admit a temporal signature. (3) Non-empty detectability band: Monitor A achieves AUC 0.52 (chance); Monitor B spans AUC 0.79-0.97 at the same 1% FPR target, and as harm is amortized over more steps Monitor A collapses to chance while Monitor B holds at AUC ~0.95. These results demonstrate a non-empty detectability band and characterize the sub-threshold sabotage frontier: distribution-shape monitors fail by construction; temporal-correlation monitors can detect but are not trivially optimal.


翻译:AI监控系统通过评估个体智能体动作来检测异常行为,但实际危害可能分散在多个看似正常的步骤中,每个步骤均低于单步报警阈值。我们构建了一种基于高斯连接函数AR(1)结构的边际保持、相关性编码分布式破坏攻击:单步监控得分边际与正常行为完全一致,因此均值、最大值、top-k尾部和阈值监控器(A类监控器)在设计上失效,而危害被编码在时间相关性结构中。我们围绕三个审稿人设定的关卡组织本文:(1)可实现性关卡:在高达3.0的所有测试危害水平下,隐蔽攻击与正常行为的KS距离仅为0.013(实际为零),证实危害与单步边际完全解耦,且可实现性不受危害限制;(2)A类与B类监控器协调:我们形式上证明,针对A类监控器得分边际构建的攻击在采用不同得分的B类监控器(相关性/序列类:CUSUM、SPRT、HMM-LR、游程检验、自相关、窗口逻辑)下仍保持边际不变,并将最坏情况声明限定于存在时间特征痕迹的得分函数;(3)非空可检测带:A类监控器AUC为0.52(随机水平);在相同1%假阳性率目标下,B类监控器AUC范围为0.79-0.97,且当危害被分摊至更多步骤时,A类监控器降至随机水平,而B类监控器保持AUC约0.95。这些结果证明存在非空可检测带,并刻画了亚阈值破坏前沿:分布形态监控器因设计限制而失效;时间相关性监控器虽能检测但并非天然最优。

0
下载
关闭预览

相关内容

《利用 LLM 进行高级持续性威胁 (APT) 检测和智能解释》
专知会员服务
24+阅读 · 2025年2月14日
《基于深度学习的实时武器检测系统》
专知会员服务
34+阅读 · 2024年1月22日
可解释人工智能中的对抗攻击和防御
专知会员服务
43+阅读 · 2023年6月20日
《边缘计算网络安全最佳实践概述》
专知会员服务
39+阅读 · 2022年7月6日
异常检测(Anomaly Detection)综述
极市平台
20+阅读 · 2020年10月24日
腾讯:机器学习构建通用的数据异常检测平台
全球人工智能
11+阅读 · 2018年5月1日
边缘计算应用:传感数据异常实时检测算法
计算机研究与发展
11+阅读 · 2018年4月10日
实战|手把手教你实现图象边缘检测!
全球人工智能
10+阅读 · 2018年1月19日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
VIP会员
最新内容
边缘计算的军事应用
专知会员服务
4+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
8+阅读 · 8月8日
《多域冲突比较支持模型》60页
专知会员服务
13+阅读 · 8月7日
相关基金
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员