Recently, Large Language Models (LLMs) have made significant advancements and are now widely used across various domains. Unfortunately, there has been a rising concern that LLMs can be misused to generate harmful or malicious content. Though a line of research has focused on aligning LLMs with human values and preventing them from producing inappropriate content, such alignments are usually vulnerable and can be bypassed by alignment-breaking attacks via adversarially optimized or handcrafted jailbreaking prompts. In this work, we introduce a Robustly Aligned LLM (RA-LLM) to defend against potential alignment-breaking attacks. RA-LLM can be directly constructed upon an existing aligned LLM with a robust alignment checking function, without requiring any expensive retraining or fine-tuning process of the original LLM. Furthermore, we also provide a theoretical analysis for RA-LLM to verify its effectiveness in defending against alignment-breaking attacks. Through real-world experiments on open-source large language models, we demonstrate that RA-LLM can successfully defend against both state-of-the-art adversarial prompts and popular handcrafted jailbreaking prompts by reducing their attack success rates from nearly 100% to around 10% or less.
翻译:近期,大语言模型取得了显著进展,并广泛应用于各个领域。然而,令人担忧的是,大语言模型可能被滥用于生成有害或恶意内容。尽管已有研究致力于使大语言模型与人类价值观对齐,并防止其生成不当内容,但这种对齐通常较为脆弱,容易受到通过对抗优化或手工设计的越狱提示发起的抗对齐破坏攻击。本文提出了一种鲁棒对齐的大语言模型,以防御潜在的对齐破坏攻击。该模型可直接基于现有已对齐的大语言模型构建,并配备鲁棒对齐检查功能,无需对原始模型进行昂贵的重训练或微调。此外,我们还提供了理论分析以验证其在防御对齐破坏攻击中的有效性。通过在开源大语言模型上的真实实验,我们证明该方法能够成功防御最先进的对抗提示和流行的手工设计越狱提示,将其攻击成功率从近100%降至约10%或更低。