Prompt Tuning is a popular parameter-efficient finetuning method for pre-trained large language models (PLMs). Recently, based on experiments with RoBERTa, it has been suggested that Prompt Tuning activates specific neurons in the transformer's feed-forward networks, that are highly predictive and selective for the given task. In this paper, we study the robustness of Prompt Tuning in relation to these "skill neurons", using RoBERTa and T5. We show that prompts tuned for a specific task are transferable to tasks of the same type but are not very robust to adversarial data, with higher robustness for T5 than RoBERTa. At the same time, we replicate the existence of skill neurons in RoBERTa and further show that skill neurons also seem to exist in T5. Interestingly, the skill neurons of T5 determined on non-adversarial data are also among the most predictive neurons on the adversarial data, which is not the case for RoBERTa. We conclude that higher adversarial robustness may be related to a model's ability to activate the relevant skill neurons on adversarial data.
翻译:提示调优是一种针对预训练大语言模型的流行参数高效微调方法。近期,基于RoBERTa的实验表明,提示调优会激活Transformer前馈网络中特定神经元,这些神经元对给定任务具有高度预测性和选择性。本文以RoBERTa和T5为研究对象,探讨与这些"技能神经元"相关的提示调优鲁棒性。我们发现:针对特定任务调优的提示可迁移至同类型任务,但对对抗数据的鲁棒性较弱,且T5的鲁棒性高于RoBERTa。与此同时,我们验证了RoBERTa中技能神经元的存在,并进一步证明T5中也存在类似神经元。值得注意的是,在非对抗数据上确定的T5技能神经元同样是预测对抗数据的核心神经元,而RoBERTa则不具备这一特性。我们得出结论:对抗鲁棒性的提升可能与模型在对抗数据上激活相关技能神经元的能力有关。