Safety, security, and compliance are essential requirements when aligning large language models (LLMs). However, many seemingly aligned LLMs are soon shown to be susceptible to jailbreak attacks. These attacks aim to circumvent the models' safety guardrails and security mechanisms by introducing jailbreak prompts into malicious queries. In response to these challenges, this paper introduces Defensive Prompt Patch (DPP), a novel prompt-based defense mechanism specifically designed to protect LLMs against such sophisticated jailbreak strategies. Unlike previous approaches, which have often compromised the utility of the model for the sake of safety, DPP is designed to achieve a minimal Attack Success Rate (ASR) while preserving the high utility of LLMs. Our method uses strategically designed interpretable suffix prompts that effectively thwart a wide range of standard and adaptive jailbreak techniques. Empirical results conducted on LLAMA-2-7B-Chat and Mistral-7B-Instruct-v0.2 models demonstrate the robustness and adaptability of DPP, showing significant reductions in ASR with negligible impact on utility. Our approach not only outperforms existing defense strategies in balancing safety and functionality, but also provides a scalable and interpretable solution applicable to various LLM platforms.
翻译:安全、可靠与合规是大语言模型对齐过程中的核心要求。然而,许多看似已对齐的模型很快被发现容易受到越狱攻击。这类攻击通过在恶意查询中植入越狱提示,旨在绕过模型的安全防护机制。为应对这些挑战,本文提出防御性提示补丁,这是一种新颖的基于提示的防御机制,专门设计用于保护大语言模型抵御此类复杂越狱策略。与以往常以牺牲模型实用性换取安全性的方法不同,DPP旨在实现最低的攻击成功率,同时保持大语言模型的高实用性。我们的方法采用经策略设计的可解释后缀提示,能有效抵御各类标准及自适应越狱技术。基于LLAMA-2-7B-Chat和Mistral-7B-Instruct-v0.2模型的实证结果表明,DPP具有显著的鲁棒性与适应性,在几乎不影响实用性的前提下显著降低了攻击成功率。该方法不仅在平衡安全性与功能性方面优于现有防御策略,更为各类大语言模型平台提供了可扩展且可解释的解决方案。