Adversarial training (AT) is an effective defense for large language models (LLMs) against jailbreak attacks, but performing AT on LLMs is costly. To improve the efficiency of AT for LLMs, recent studies propose continuous AT (CAT) that searches for adversarial inputs within the continuous embedding space of LLMs during AT. While CAT has achieved empirical success, its underlying mechanism, i.e., why adversarial perturbations in the embedding space can help LLMs defend against jailbreak prompts synthesized in the input token space, remains unknown. This paper presents the first theoretical analysis of CAT on LLMs based on in-context learning (ICL) theory. For linear transformers trained with adversarial examples from the embedding space on in-context linear regression tasks, we prove a robust generalization bound that has a negative correlation with the perturbation radius in the embedding space. This clearly explains why CAT can defend against jailbreak prompts from the LLM's token space. Further, the robust bound shows that the robustness of an adversarially trained LLM is closely related to the singular values of its embedding matrix. Based on this, we propose to improve LLM CAT by introducing an additional regularization term, which depends on singular values of the LLM's embedding matrix, into the objective function of CAT. Experiments on real-world LLMs demonstrate that our method can help LLMs achieve a better jailbreak robustness-utility tradeoff. The code is available at https://github.com/fshp971/continuous-adv-icl.


翻译:对抗训练(AT)是大语言模型(LLMs)抵御越狱攻击的有效防御手段,但直接对LLMs进行AT成本高昂。为提高LLMs对抗训练的效率,近期研究提出了连续对抗训练(CAT),该方法在对抗训练过程中于LLMs的连续嵌入空间中搜索对抗输入。尽管CAT已取得经验性成功,但其内在机理——为何嵌入空间的对抗扰动能帮助LLMs抵御输入词元空间合成的越狱提示——仍不明确。本文基于上下文学习(ICL)理论,首次对LLMs的CAT进行了理论分析。针对在上下文线性回归任务中使用嵌入空间对抗样本训练的线性Transformer,我们证明了一个鲁棒泛化界,该界与嵌入空间中的扰动半径呈负相关。这清晰解释了CAT为何能防御来自LLMs词元空间的越狱提示。此外,该鲁棒泛化界表明,经过对抗训练的LLMs的鲁棒性与其嵌入矩阵的奇异值密切相关。基于此,我们提出通过在CAT目标函数中引入一个依赖于LLMs嵌入矩阵奇异值的额外正则化项来改进CAT。在真实LLMs上的实验表明,我们的方法能帮助LLMs实现更好的越狱鲁棒性-效用权衡。代码已开源:https://github.com/fshp971/continuous-adv-icl。

0
下载
关闭预览

相关内容

大语言模型持续学习:方法、挑战与机遇
专知会员服务
22+阅读 · 3月16日
Llama-3-SynE:实现有效且高效的大语言模型持续预训练
专知会员服务
36+阅读 · 2024年7月30日
《大型语言模型持续学习》综述
专知会员服务
94+阅读 · 2024年4月26日
大型语言模型增强强化学习综述:概念、分类和方法
专知会员服务
57+阅读 · 2024年4月4日
基于模型的强化学习综述
专知
42+阅读 · 2022年7月13日
国家自然科学基金
44+阅读 · 2015年12月31日
国家自然科学基金
41+阅读 · 2015年12月31日
国家自然科学基金
20+阅读 · 2015年12月31日
国家自然科学基金
24+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
VIP会员
相关主题
最新内容
反制无人机:乌克兰提供的五点启示
专知会员服务
7+阅读 · 9月23日
《各指挥层级均亟需红队能力》报告
专知会员服务
7+阅读 · 9月23日
《航电任务系统框架(FAMOS)》50页报告
专知会员服务
4+阅读 · 9月22日
《对抗行动中的人工智能与自主性》智库报告
专知会员服务
8+阅读 · 9月22日
《从数据到胜利:战争中的分析优势之争》
专知会员服务
10+阅读 · 9月22日
战争不仅需要机器人:人类仍不可或缺
专知会员服务
6+阅读 · 9月21日
《描绘美国防部创新基础设施的未来蓝图》100页
专知会员服务
10+阅读 · 9月21日
相关基金
国家自然科学基金
44+阅读 · 2015年12月31日
国家自然科学基金
41+阅读 · 2015年12月31日
国家自然科学基金
20+阅读 · 2015年12月31日
国家自然科学基金
24+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员