LLM-based guardrails have emerged as a highly effective defense against prompt injection and jailbreak attacks in autonomous agents. However, we reveal that the very reasoning and task-following capabilities enabling this protection introduce a novel vulnerability: attackers can inject crafted data to trap the guardrail in extended reasoning loops, effectuating a systematic denial-of-service (DoS) attack. To systematically expose this threat, we design a beam-search optimization framework that crafts natural-language payloads to maximize guardrail reasoning length, utilizing an LLM proposer guided by a strategy bank. Based on the observation of guardrail's schema-following nature, we also provide another attack framework driven by mechanism-aware structural mutations with less computational load. The attack efficacy is systematically evaluated in two parts. First, in standalone evaluations, the attack generalizes across diverse guardrail architectures, safety templates, and agent benchmarks. Payloads optimized on a single open-source surrogate successfully transfer to eight leading model backbones (e.g., Claude, GPT, Gemini, DeepSeek, and Qwen), achieving a 13--63$\times$ token amplification. Second, in end-to-end real-world agent deployments (web, desktop, code, and multi-agent systems), the attack reveals up to a 148$\times$ latency amplification. We show that a single poisoned document can saturate shared guardrail infrastructures, effectively starving co-located agents and paralyzing the entire system. By uncovering this availability flaw, our work underscores the urgent need to develop cost-bounded, reasoning-robust guardrails.


翻译:基于LLM的护栏已成为自主智能体中抵御提示注入和越狱攻击的高效防御手段。然而,我们发现,正是这种实现保护的推理和任务遵循能力引入了新的脆弱性:攻击者可注入精心构造的数据,使护栏陷入扩展推理循环,从而实施系统性拒绝服务(DoS)攻击。为系统揭示这一威胁,我们设计了一种束搜索优化框架,利用由策略库引导的LLM提议器生成自然语言载荷,最大化护栏推理长度。基于护栏遵循架构模式的观察,我们还提供了另一种由机制感知的结构突变驱动的攻击框架,其计算开销更低。攻击效能通过两个部分系统评估。首先,在独立评估中,该攻击可泛化至多种护栏架构、安全模板及智能体基准测试。在单一开源替代模型上优化的载荷成功迁移至八个主流模型主干(如Claude、GPT、Gemini、DeepSeek和Qwen),实现了13-63倍的令牌放大。其次,在端到端真实世界智能体部署(包括Web、桌面、代码及多智能体系统)中,该攻击显示出高达148倍的延迟放大。我们证明,单个有毒文档可饱和共享护栏基础设施,有效饿死共置智能体并瘫痪整个系统。通过揭示这一可用性缺陷,我们的工作强调了开发成本受限、推理鲁棒护栏的迫切需求。

0
下载
关闭预览

相关内容

综述 | Always-On Agents:LLM智能体持久状态与治理
专知会员服务
22+阅读 · 6月30日
智能体安全综述:应用、威胁与防御
专知会员服务
45+阅读 · 2025年10月12日
LLM/智能体作为数据分析师:综述
专知会员服务
38+阅读 · 2025年9月30日
可信赖LLM智能体的研究综述:威胁与应对措施
专知会员服务
36+阅读 · 2025年3月17日
国防中的LLM:五角大楼的机遇与挑战
专知会员服务
44+阅读 · 2024年3月5日
ICLR2019 图上的对抗攻击
图与推荐
17+阅读 · 2020年3月15日
探秘各种主流周界安防技术产品
未来产业促进会
12+阅读 · 2018年11月16日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
19+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2012年12月31日
VIP会员
最新内容
边缘计算的军事应用
专知会员服务
6+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
8+阅读 · 8月8日
《多域冲突比较支持模型》60页
专知会员服务
13+阅读 · 8月7日
相关基金
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
19+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员