Caution: This paper includes offensive words that could potentially cause unpleasantness. Language models (LMs) are vulnerable to exploitation for adversarial misuse. Training LMs for safety alignment is extensive and makes it hard to respond to fast-developing attacks immediately, such as jailbreaks. We propose self-refine with formatting that achieves outstanding safety even in non-safety-aligned LMs and evaluate our method alongside several defense baselines, demonstrating that it is the safest training-free method against jailbreak attacks. Additionally, we proposed a formatting method that improves the efficiency of the self-refine process while reducing attack success rates in fewer iterations. We've also observed that non-safety-aligned LMs outperform safety-aligned LMs in safety tasks by giving more helpful and safe responses. In conclusion, our findings can achieve less safety risk with fewer computational costs, allowing non-safety LM to be easily utilized in real-world service.
翻译:注意:本文包含可能引发不适的冒犯性词汇。语言模型(LM)易被利用进行对抗性滥用。为保障安全性而对LM进行对齐训练的过程极为繁琐,导致其难以快速应对诸如越狱等迅速演变的攻击。我们提出了一种结合格式化的自我精炼方法,即使未经过安全对齐的LM也能展现出卓越的安全性。在与多种防御基线方法的对比评估中,该方法被证明是当前最安全的免训练越狱攻击防御手段。此外,我们提出的一种格式化方法能在减少迭代次数的同时提升自我精炼流程的效率,并降低攻击成功率。我们还发现,在安全任务中,未经过安全对齐的LM相较于安全对齐的LM,能提供更具实用性且安全性的响应。综上,本研究可在降低计算成本的同时减少安全风险,使未经过安全对齐的LM能便捷地应用于实际服务场景。