LLM-based guardrails have emerged as a highly effective defense against prompt injection and jailbreak attacks in autonomous agents. However, we reveal that the very reasoning and task-following capabilities enabling this protection introduce a novel vulnerability: attackers can inject crafted data to trap the guardrail in extended reasoning loops, effectuating a systematic denial-of-service (DoS) attack. To systematically expose this threat, we design a beam-search optimization framework that crafts natural-language payloads to maximize guardrail reasoning length, utilizing an LLM proposer guided by a strategy bank. Based on the observation of guardrail's schema-following nature, we also provide another attack framework driven by mechanism-aware structural mutations with less computational load. The attack efficacy is systematically evaluated in two parts. First, in standalone evaluations, the attack generalizes across diverse guardrail architectures, safety templates, and agent benchmarks. Payloads optimized on a single open-source surrogate successfully transfer to eight leading model backbones (e.g., Claude, GPT, Gemini, DeepSeek, and Qwen), achieving a 13--63$\times$ token amplification. Second, in end-to-end real-world agent deployments (web, desktop, code, and multi-agent systems), the attack reveals up to a 148$\times$ latency amplification. We show that a single poisoned document can saturate shared guardrail infrastructures, effectively starving co-located agents and paralyzing the entire system. By uncovering this availability flaw, our work underscores the urgent need to develop cost-bounded, reasoning-robust guardrails.
翻译:基于LLM的护栏已成为自主智能体中抵御提示注入和越狱攻击的高效防御手段。然而,我们发现,正是这种实现保护的推理和任务遵循能力引入了新的脆弱性:攻击者可注入精心构造的数据,使护栏陷入扩展推理循环,从而实施系统性拒绝服务(DoS)攻击。为系统揭示这一威胁,我们设计了一种束搜索优化框架,利用由策略库引导的LLM提议器生成自然语言载荷,最大化护栏推理长度。基于护栏遵循架构模式的观察,我们还提供了另一种由机制感知的结构突变驱动的攻击框架,其计算开销更低。攻击效能通过两个部分系统评估。首先,在独立评估中,该攻击可泛化至多种护栏架构、安全模板及智能体基准测试。在单一开源替代模型上优化的载荷成功迁移至八个主流模型主干(如Claude、GPT、Gemini、DeepSeek和Qwen),实现了13-63倍的令牌放大。其次,在端到端真实世界智能体部署(包括Web、桌面、代码及多智能体系统)中,该攻击显示出高达148倍的延迟放大。我们证明,单个有毒文档可饱和共享护栏基础设施,有效饿死共置智能体并瘫痪整个系统。通过揭示这一可用性缺陷,我们的工作强调了开发成本受限、推理鲁棒护栏的迫切需求。