Large Language Models (LLMs) are typically harmless but remain vulnerable to carefully crafted prompts known as ``jailbreaks'', which can bypass protective measures and induce harmful behavior. Recent advancements in LLMs have incorporated moderation guardrails that can filter outputs, which trigger processing errors for certain malicious questions. Existing red-teaming benchmarks often neglect to include questions that trigger moderation guardrails, making it difficult to evaluate jailbreak effectiveness. To address this issue, we introduce JAMBench, a harmful behavior benchmark designed to trigger and evaluate moderation guardrails. JAMBench involves 160 manually crafted instructions covering four major risk categories at multiple severity levels. Furthermore, we propose a jailbreak method, JAM (Jailbreak Against Moderation), designed to attack moderation guardrails using jailbreak prefixes to bypass input-level filters and a fine-tuned shadow model functionally equivalent to the guardrail model to generate cipher characters to bypass output-level filters. Our extensive experiments on four LLMs demonstrate that JAM achieves higher jailbreak success ($\sim$ $\times$ 19.88) and lower filtered-out rates ($\sim$ $\times$ 1/6) than baselines.
翻译:大语言模型(LLMs)通常是无害的,但仍易受精心设计的提示(即“越狱”攻击)影响,这些提示可能绕过保护措施并诱导有害行为。近期LLMs的进展已引入能够过滤输出的内容审核防护机制,该机制会对特定恶意问题触发处理错误。现有的红队测试基准往往未包含能触发审核防护机制的问题,导致难以评估越狱攻击的有效性。为解决此问题,我们提出了JAMBench——一个专门用于触发并评估内容审核防护机制的有害行为基准。JAMBench包含160条人工构建的指令,覆盖四大风险类别及多个严重等级。此外,我们提出了一种名为JAM(针对内容审核的越狱攻击)的越狱方法,该方法通过越狱前缀绕过输入级过滤器,并利用功能等效于防护模型的微调影子模型生成密文字符以绕过输出级过滤器。我们在四个LLM上的大量实验表明,与基线方法相比,JAM实现了更高的越狱成功率(约19.88倍)和更低的过滤率(约1/6)。