Large Language Models (LLMs) such as GPT and Llama2 are increasingly adopted in many safety-critical applications. Their security is thus essential. Even with considerable efforts spent on reinforcement learning from human feedback (RLHF), recent studies have shown that LLMs are still subject to attacks such as adversarial perturbation and Trojan attacks. Further research is thus needed to evaluate their security and/or understand the lack of it. In this work, we propose a framework for conducting light-weight causality-analysis of LLMs at the token, layer, and neuron level. We applied our framework to open-source LLMs such as Llama2 and Vicuna and had multiple interesting discoveries. Based on a layer-level causality analysis, we show that RLHF has the effect of overfitting a model to harmful prompts. It implies that such security can be easily overcome by `unusual' harmful prompts. As evidence, we propose an adversarial perturbation method that achieves 100\% attack success rate on the red-teaming tasks of the Trojan Detection Competition 2023. Furthermore, we show the existence of one mysterious neuron in both Llama2 and Vicuna that has an unreasonably high causal effect on the output. While we are uncertain on why such a neuron exists, we show that it is possible to conduct a ``Trojan'' attack targeting that particular neuron to completely cripple the LLM, i.e., we can generate transferable suffixes to prompts that frequently make the LLM produce meaningless responses.
翻译:诸如GPT和Llama2等大型语言模型(LLMs)正被越来越多地应用于众多安全关键型场景,因此其安全性至关重要。尽管在基于人类反馈的强化学习(RLHF)方面投入了大量精力,但近期研究表明,LLMs仍易受到对抗性扰动及木马攻击等威胁。因此,需进一步研究以评估其安全性及/或理解其安全缺失的根源。本文提出了一种轻量级因果分析框架,可在词元、层及神经元级别对LLMs进行分析。我们将该框架应用于Llama2和Vicuna等开源LLMs,并获得多项有趣发现。基于层级因果分析,我们发现RLHF会导致模型对有害提示产生过拟合效应。这意味着此类安全性可轻易被"非常规"有害提示所突破。作为佐证,我们提出一种对抗性扰动方法,在2023年木马检测竞赛的红队任务中实现了100%的攻击成功率。此外,我们揭示了Llama2与Vicuna中存在一个神秘神经元,其对输出具有异常高的因果效应。尽管我们尚不明确此类神经元存在的原因,但研究表明,针对该特定神经元实施"木马"攻击可完全瘫痪LLM——即能够生成可迁移的后缀附加于提示后,导致LLM频繁输出无意义响应。