Ensuring the safety alignment of Large Language Models (LLMs) is crucial to generating responses consistent with human values. Despite their ability to recognize and avoid harmful queries, LLMs are vulnerable to "jailbreaking" attacks, where carefully crafted prompts elicit them to produce toxic content. One category of jailbreak attacks is reformulating the task as adversarial attacks by eliciting the LLM to generate an affirmative response. However, the typical attack in this category GCG has very limited attack success rate. In this study, to better study the jailbreak attack, we introduce the DSN (Don't Say No) attack, which prompts LLMs to not only generate affirmative responses but also novelly enhance the objective to suppress refusals. In addition, another challenge lies in jailbreak attacks is the evaluation, as it is difficult to directly and accurately assess the harmfulness of the attack. The existing evaluation such as refusal keyword matching has its own limitation as it reveals numerous false positive and false negative instances. To overcome this challenge, we propose an ensemble evaluation pipeline incorporating Natural Language Inference (NLI) contradiction assessment and two external LLM evaluators. Extensive experiments demonstrate the potency of the DSN and the effectiveness of ensemble evaluation compared to baseline methods.
翻译:确保大语言模型(LLM)的安全对齐对于生成符合人类价值观的回复至关重要。尽管LLM能够识别并避免有害查询,但它们容易受到"越狱"攻击的威胁——精心设计的提示词可诱使其产生有毒内容。越狱攻击的其中一种策略是将任务重构为对抗性攻击,通过诱导LLM生成肯定性回应。然而,该类别的典型攻击GCG的攻击成功率非常有限。在本研究中,为更好研究越狱攻击,我们提出了DSN(不要说"不")攻击方法,该攻击不仅促使LLM生成肯定性回应,更创新性地增强了抑制拒绝机制的目标。此外,越狱攻击的另一挑战在于评估,因为难以直接准确衡量攻击的危害性。现有的拒绝关键词匹配评估方法存在局限性,会产生大量假阳性和假阴性案例。为克服这一挑战,我们提出了集成评估流水线,融合了自然语言推理(NLI)矛盾评估和两个外部LLM评估器。大量实验证明,与基线方法相比,DSN攻击具有显著效力,集成评估方法也展现出更优效果。