Applicating third-party data and models has become a new paradigm for language modeling in NLP, which also introduces some potential security vulnerabilities because attackers can manipulate the training process and data source. In this case, backdoor attacks can induce the model to exhibit expected behaviors through specific triggers and have little inferior influence on primitive tasks. Hence, it could have dire consequences, especially considering that the backdoor attack surfaces are broad. However, there is still no systematic and comprehensive review to reflect the security challenges, attacker's capabilities, and purposes according to the attack surface. Moreover, there is a shortage of analysis and comparison of the diverse emerging backdoor countermeasures in this context. In this paper, we conduct a timely review of backdoor attacks and countermeasures to sound the red alarm for the NLP security community. According to the affected stage of the machine learning pipeline, the attack surfaces are recognized to be wide and then formalized into three categorizations: attacking pre-trained model with fine-tuning (APMF) or parameter-efficient tuning (APMP), and attacking final model with training (AFMT). Thus, attacks under each categorization are combed. The countermeasures are categorized into two general classes: sample inspection and model inspection. Overall, the research on the defense side is far behind the attack side, and there is no single defense that can prevent all types of backdoor attacks. An attacker can intelligently bypass existing defenses with a more invisible attack. Drawing the insights from the systematic review, we also present crucial areas for future research on the backdoor, such as empirical security evaluations on large language models, and in particular, more efficient and practical countermeasures are solicited.
翻译:应用第三方数据和模型已成为自然语言处理中语言建模的新范式,但这也引入了潜在安全漏洞,因为攻击者可操纵训练过程与数据源。在此背景下,后门攻击能够通过特定触发器诱导模型产生预期行为,同时对原始任务影响甚微。鉴于后门攻击面广阔,其可能造成严重后果。然而,目前仍缺乏系统性综述来反映不同攻击面所对应的安全挑战、攻击者能力及意图,且对该背景下新兴后门防御对策的分析与比较存在空白。本文对后门攻击及防御对策进行了及时综述,为自然语言处理安全社区敲响警钟。根据机器学习流程中受影响的阶段,攻击面被识别为广泛存在,并形式化为三类:攻击微调预训练模型、攻击参数高效微调预训练模型及攻击训练最终模型。据此梳理了各类攻击方法,并将防御对策归纳为样本检测与模型检测两大类。总体而言,防御领域研究远落后于攻击领域,尚无单一防御机制可抵御所有类型后门攻击。攻击者可通过更隐蔽的攻击手段智能规避现有防御。基于系统性综述的洞察,本文还提出了后门技术未来研究的关键方向,例如针对大语言模型开展实证安全评估,尤其是需要开发更高效、实用的防御对策。