Adversaries can embed backdoors in deep learning models by introducing backdoor poison samples into training datasets. In this work, we investigate how to detect such poison samples to mitigate the threat of backdoor attacks. First, we uncover a post-hoc workflow underlying most prior work, where defenders passively allow the attack to proceed and then leverage the characteristics of the post-attacked model to uncover poison samples. We reveal that this workflow does not fully exploit defenders' capabilities, and defense pipelines built on it are prone to failure or performance degradation in many scenarios. Second, we suggest a paradigm shift by promoting a proactive mindset in which defenders engage proactively with the entire model training and poison detection pipeline, directly enforcing and magnifying distinctive characteristics of the post-attacked model to facilitate poison detection. Based on this, we formulate a unified framework and provide practical insights on designing detection pipelines that are more robust and generalizable. Third, we introduce the technique of Confusion Training (CT) as a concrete instantiation of our framework. CT applies an additional poisoning attack to the already poisoned dataset, actively decoupling benign correlation while exposing backdoor patterns to detection. Empirical evaluations on 4 datasets and 14 types of attacks validate the superiority of CT over 14 baseline defenses.
翻译:攻击者通过在训练数据集中引入后门投毒样本,可将后门嵌入深度学习模型。本文研究了如何检测此类投毒样本,以减轻后门攻击的威胁。首先,我们揭示了大多数现有研究遵循的一种事后工作流程:防御者被动地允许攻击进行,随后利用受攻击后模型的特征来发现投毒样本。我们发现该流程未能充分利用防御者的能力,基于此构建的防御管线在多种场景下容易失效或性能下降。其次,我们倡导一种范式转变,即推广主动式思维:防御者主动介入整个模型训练和投毒检测过程,直接强化并放大受攻击后模型的独特特征,以促进投毒检测。基于此,我们构建了一个统一框架,并为设计更具鲁棒性和泛化性的检测管线提供了实践见解。第三,我们引入了混淆训练技术,作为所提框架的具体实现。混淆训练对已投毒数据集施加额外的攻击,主动解耦良性关联,同时将后门模式暴露于检测手段之下。在4个数据集和14种攻击类型上的实验验证了混淆训练相较于14种基线防御方法的优越性。