Prompt learning has been proven to be highly effective in improving pre-trained language model (PLM) adaptability, surpassing conventional fine-tuning paradigms, and showing exceptional promise in an ever-growing landscape of applications and APIs tailored for few-shot learning scenarios. Despite the growing prominence of prompt learning-based APIs, their security concerns remain underexplored. In this paper, we undertake a pioneering study on the Trojan susceptibility of prompt-learning PLM APIs. We identified several key challenges, including discrete-prompt, few-shot, and black-box settings, which limit the applicability of existing backdoor attacks. To address these challenges, we propose TrojPrompt, an automatic and black-box framework to effectively generate universal and stealthy triggers and insert Trojans into hard prompts. Specifically, we propose a universal API-driven trigger discovery algorithm for generating universal triggers for various inputs by querying victim PLM APIs using few-shot data samples. Furthermore, we introduce a novel progressive trojan poisoning algorithm designed to generate poisoned prompts that retain efficacy and transferability across a diverse range of models. Our experiments and results demonstrate TrojPrompt's capacity to effectively insert Trojans into text prompts in real-world black-box PLM APIs, while maintaining exceptional performance on clean test sets and significantly outperforming baseline models. Our work sheds light on the potential security risks in current models and offers a potential defensive approach.
翻译:提示学习已被证明在提升预训练语言模型(PLM)适应性方面极为有效,超越了传统的微调范式,并在面向少样本学习场景的日益增长的应用程序和API生态中展现出卓越前景。尽管基于提示学习的API应用日益广泛,但其安全性问题仍未得到充分探索。本文首次系统研究了提示学习PLM API的木马攻击脆弱性。我们识别出若干关键挑战,包括离散提示、少样本和黑盒设置,这些因素限制了现有后门攻击的适用性。为应对这些挑战,我们提出TrojPrompt——一个自动化的黑盒框架,能够高效生成通用且隐蔽的触发器,并将木马注入硬提示中。具体而言,我们提出一种通用API驱动的触发器发现算法,通过利用少样本数据查询受害PLM API,为各种输入生成通用触发器。此外,我们引入一种新颖的渐进式木马投毒算法,旨在生成在多种模型间保持有效性和可迁移性的投毒提示。实验结果表明,TrojPrompt能够在实际黑盒PLM API中有效将木马注入文本提示,同时在干净测试集上保持卓越性能,并显著优于基线模型。本研究揭示了当前模型中潜在的安全风险,并提供了一种可能的防御思路。