Large Language Models (LLMs), such as ChatGPT, encounter `jailbreak' challenges, wherein safeguards are circumvented to generate ethically harmful prompts. This study introduces a straightforward black-box method for efficiently crafting jailbreak prompts, addressing the significant complexity and computational costs associated with conventional methods. Our technique iteratively transforms harmful prompts into benign expressions directly utilizing the target LLM, predicated on the hypothesis that LLMs can autonomously generate expressions that evade safeguards. Through experiments conducted with ChatGPT (GPT-3.5 and GPT-4) and Gemini-Pro, our method consistently achieved an attack success rate exceeding 80% within an average of five iterations for forbidden questions and proved robust against model updates. The jailbreak prompts generated were not only naturally-worded and succinct but also challenging to defend against. These findings suggest that the creation of effective jailbreak prompts is less complex than previously believed, underscoring the heightened risk posed by black-box jailbreak attacks.
翻译:大型语言模型(LLM),例如ChatGPT,面临着“越狱”挑战,即安全防护措施被绕过以生成违背伦理的提示内容。本研究提出了一种简单的黑盒方法,用于高效构造越狱提示,解决了传统方法中显著的高复杂性与计算成本问题。我们的技术直接利用目标LLM,将有害提示迭代转化为良性表述,该方法的假设基础是LLM能自主生成可规避安全防护的表述。通过在ChatGPT(GPT-3.5与GPT-4)及Gemini-Pro上开展的实验,该方法对禁止提出的问题在平均五次迭代内始终实现了超过80%的攻击成功率,并证明其能抵御模型更新。所生成的越狱提示不仅措辞自然、简洁,而且难以防御。这些发现表明,有效越狱提示的构造难度远低于先前认知,突显了黑盒越狱攻击所带来的更高风险。