This paper proposes a jailbreaking prompt detection method for large language models (LLMs) to defend against jailbreak attacks. Although recent LLMs are equipped with built-in safeguards, it remains possible to craft jailbreaking prompts that bypass them. We argue that such jailbreaking prompts are inherently fragile, and thus introduce an embedding disruption method to re-activate the safeguards within LLMs. Unlike previous defense methods that aim to serve as standalone solutions, our approach instead cooperates with the LLM's internal defense mechanisms by re-triggering them. Moreover, through extensive analysis, we gain a comprehensive understanding of the disruption effects and develop an efficient search algorithm to identify appropriate disruptions for effective jailbreak detection. Extensive experiments demonstrate that our approach effectively defends against state-of-the-art jailbreak attacks in white-box and black-box settings, and remains robust even against adaptive attacks.
翻译:本文提出了一种针对大语言模型(LLMs)的越狱提示检测方法,以防御越狱攻击。尽管近期的大语言模型配备了内置防护机制,但攻击者仍可能构造绕过这些机制的越狱提示。我们认为这类越狱提示本质上是脆弱的,因此引入了一种嵌入扰动方法来重新激活大语言模型内部的防护机制。与以往旨在作为独立解决方案的防御方法不同,我们的方法通过重新触发大语言模型的内部防御机制与其协同工作。此外,通过广泛分析,我们深入理解了扰动效应,并开发了一种高效的搜索算法来识别合适的扰动以实现有效的越狱检测。大量实验表明,我们的方法在白盒和黑盒场景下均能有效防御最先进的越狱攻击,并且对自适应攻击具有鲁棒性。