Backdoor (Trojan) attack is a common threat to deep neural networks, where samples from one or more source classes embedded with a backdoor trigger will be misclassified to adversarial target classes. Existing methods for detecting whether a classifier is backdoor attacked are mostly designed for attacks with a single adversarial target (e.g., all-to-one attack). To the best of our knowledge, without supervision, no existing methods can effectively address the more general X2X attack with an arbitrary number of source classes, each paired with an arbitrary target class. In this paper, we propose UMD, the first Unsupervised Model Detection method that effectively detects X2X backdoor attacks via a joint inference of the adversarial (source, target) class pairs. In particular, we first define a novel transferability statistic to measure and select a subset of putative backdoor class pairs based on a proposed clustering approach. Then, these selected class pairs are jointly assessed based on an aggregation of their reverse-engineered trigger size for detection inference, using a robust and unsupervised anomaly detector we proposed. We conduct comprehensive evaluations on CIFAR-10, GTSRB, and Imagenette dataset, and show that our unsupervised UMD outperforms SOTA detectors (even with supervision) by 17%, 4%, and 8%, respectively, in terms of the detection accuracy against diverse X2X attacks. We also show the strong detection performance of UMD against several strong adaptive attacks.
翻译:后门(木马)攻击是对深度神经网络的常见威胁,当嵌入后门触发器的样本来自一个或多个源类别时,会被错误分类至对抗目标类别。现有检测分类器是否遭受后门攻击的方法大多针对单一对抗目标(例如全对一攻击)设计。据我们所知,在没有监督的情况下,现有方法无法有效应对更具普适性的X2X攻击——其源类别数量任意,且每个源类别与任意目标类别配对。本文提出UMD,这是首个通过联合推断对抗(源、目标)类别对来有效检测X2X后门攻击的无监督模型检测方法。具体而言,我们首先定义一种新颖的可迁移性统计量,基于所提出的聚类方法测量并选择一组疑似后门类别对。随后,利用我们提出的鲁棒无监督异常检测器,通过聚合这些选定类别对的反向工程触发器尺寸进行联合评估,以完成检测推断。我们在CIFAR-10、GTSRB和Imagenette数据集上开展了全面评估,结果表明:针对多种X2X攻击,我们的无监督UMD在检测准确率上分别超越现有最先进的检测器(包括有监督方法)17%、4%和8%。此外,我们还展示了UMD在面对多种强适应性攻击时仍具有强大的检测性能。