The prompt-based learning paradigm, which bridges the gap between pre-training and fine-tuning, achieves state-of-the-art performance on several NLP tasks, particularly in few-shot settings. Despite being widely applied, prompt-based learning is vulnerable to backdoor attacks. Textual backdoor attacks are designed to introduce targeted vulnerabilities into models by poisoning a subset of training samples through trigger injection and label modification. However, they suffer from flaws such as abnormal natural language expressions resulting from the trigger and incorrect labeling of poisoned samples. In this study, we propose ProAttack, a novel and efficient method for performing clean-label backdoor attacks based on the prompt, which uses the prompt itself as a trigger. Our method does not require external triggers and ensures correct labeling of poisoned samples, improving the stealthy nature of the backdoor attack. With extensive experiments on rich-resource and few-shot text classification tasks, we empirically validate ProAttack's competitive performance in textual backdoor attacks. Notably, in the rich-resource setting, ProAttack achieves state-of-the-art attack success rates in the clean-label backdoor attack benchmark without external triggers.
翻译:基于提示的学习范式弥合了预训练与微调之间的差距,在多项自然语言处理任务中(尤其在小样本场景下)达到了最先进的性能。尽管被广泛应用,基于提示的学习易受后门攻击。文本后门攻击通过触发词注入和标签修改来污染部分训练样本,从而向模型中引入特定漏洞。然而,这类攻击存在缺陷,例如触发词导致的异常自然语言表达以及污染样本的错误标注。在本研究中,我们提出ProAttack——一种基于提示的、高效且新颖的干净标签后门攻击方法,该方法直接利用提示本身作为触发器。我们的方法无需外部触发词,并确保污染样本的正确标注,从而提升了后门攻击的隐蔽性。通过在丰富资源和小样本文本分类任务上的大量实验,我们实证验证了ProAttack在文本后门攻击中的竞争性能。值得注意的是,在丰富资源场景下,ProAttack无需外部触发词即可在干净标签后门攻击基准测试中达到最先进的攻击成功率。