Large Language Models (LLMs), such as ChatGPT and GPT-4, are designed to provide useful and safe responses. However, adversarial prompts known as 'jailbreaks' can circumvent safeguards, leading LLMs to generate harmful content. Exploring jailbreak prompts can help to better reveal the weaknesses of LLMs and further steer us to secure them. Unfortunately, existing jailbreak methods either suffer from intricate manual design or require optimization on another white-box model, compromising generalization or jailbreak efficiency. In this paper, we generalize jailbreak prompt attacks into two aspects: (1) Prompt Rewriting and (2) Scenario Nesting. Based on this, we propose ReNeLLM, an automatic framework that leverages LLMs themselves to generate effective jailbreak prompts. Extensive experiments demonstrate that ReNeLLM significantly improves the attack success rate while greatly reducing the time cost compared to existing baselines. Our study also reveals the inadequacy of current defense methods in safeguarding LLMs. Finally, we offer detailed analysis and discussion from the perspective of prompt execution priority on the failure of LLMs' defense. We hope that our research can catalyze both the academic community and LLMs vendors towards the provision of safer and more regulated Large Language Models.
翻译:大型语言模型(LLMs),例如ChatGPT和GPT-4,旨在提供有用且安全的回复。然而,被称为“越狱”的对抗性提示可能绕过安全措施,导致LLMs生成有害内容。探索越狱提示有助于更好地揭示LLMs的弱点,并进一步引导我们确保其安全性。不幸的是,现有的越狱方法要么依赖复杂的手工设计,要么需要在另一个白盒模型上进行优化,从而牺牲了通用性或越狱效率。在本文中,我们将越狱提示攻击概括为两个方面:(1)提示重写和(2)场景嵌套。基于此,我们提出了ReNeLLM,一种利用LLMs自身自动生成有效越狱提示的框架。大量实验表明,与现有基线相比,ReNeLLM显著提升了攻击成功率,同时大幅降低了时间成本。我们的研究还揭示了当前防御方法在保护LLMs方面的不足。最后,我们从提示执行优先级的角度,对LLMs防御失败的原因进行了详细分析和讨论。我们希望我们的研究能推动学术界和LLMs供应商提供更安全、更规范的大型语言模型。