The jailbreak attack can bypass the safety measures of a Large Language Model (LLM), generating harmful content. This misuse of LLM has led to negative societal consequences. Currently, there are two main approaches to address jailbreak attacks: safety training and safeguards. Safety training focuses on further training LLM to enhance its safety. On the other hand, safeguards involve implementing external models or filters to prevent harmful outputs. However, safety training has constraints in its ability to adapt to new attack types and often leads to a drop in model performance. Safeguards have proven to be of limited help. To tackle these issues, we propose a novel approach called Self-Guard, which combines the strengths of both safety methods. Self-Guard includes two stages. In the first stage, we enhance the model's ability to assess harmful content, and in the second stage, we instruct the model to consistently perform harmful content detection on its own responses. The experiment has demonstrated that Self-Guard is robust against jailbreak attacks. In the bad case analysis, we find that LLM occasionally provides harmless responses to harmful queries. Additionally, we evaluated the general capabilities of the LLM before and after safety training, providing evidence that Self-Guard does not result in the LLM's performance degradation. In sensitivity tests, Self-Guard not only avoids inducing over-sensitivity in LLM but also can even mitigate this issue.
翻译:越狱攻击能够绕过大型语言模型(LLM)的安全措施,生成有害内容。这种对LLM的滥用已导致负面的社会后果。目前,应对越狱攻击主要有两种方法:安全训练和安全防护。安全训练侧重于通过进一步训练LLM来增强其安全性,而安全防护则涉及实施外部模型或过滤器来防止有害输出。然而,安全训练在适应新型攻击方面存在局限性,且常导致模型性能下降;安全防护则被证明帮助有限。为解决这些问题,我们提出了一种名为Self-Guard的新方法,它融合了两种安全方法的优势。Self-Guard包括两个阶段:第一阶段增强模型评估有害内容的能力,第二阶段引导模型持续对其自身响应进行有害内容检测。实验表明,Self-Guard对越狱攻击具有鲁棒性。在不良案例分析中,我们发现LLM偶尔会对有害查询给出无害响应。此外,我们评估了安全训练前后LLM的通用能力,为Self-Guard不会导致LLM性能退化提供了证据。在敏感性测试中,Self-Guard不仅能避免诱导LLM过度敏感,甚至还能缓解这一问题。