To gather a significant quantity of annotated training data for high-performance image classification models, numerous companies opt to enlist third-party providers to label their unlabeled data. This practice is widely regarded as secure, even in cases where some annotated errors occur, as the impact of these minor inaccuracies on the final performance of the models is negligible and existing backdoor attacks require attacker's ability to poison the training images. Nevertheless, in this paper, we propose clean-image backdoor attacks which uncover that backdoors can still be injected via a fraction of incorrect labels without modifying the training images. Specifically, in our attacks, the attacker first seeks a trigger feature to divide the training images into two parts: those with the feature and those without it. Subsequently, the attacker falsifies the labels of the former part to a backdoor class. The backdoor will be finally implanted into the target model after it is trained on the poisoned data. During the inference phase, the attacker can activate the backdoor in two ways: slightly modifying the input image to obtain the trigger feature, or taking an image that naturally has the trigger feature as input. We conduct extensive experiments to demonstrate the effectiveness and practicality of our attacks. According to the experimental results, we conclude that our attacks seriously jeopardize the fairness and robustness of image classification models, and it is necessary to be vigilant about the incorrect labels in outsourced labeling.
翻译:为了收集大量标注训练数据以训练高性能图像分类模型,许多公司选择委托第三方供应商标注其未标注数据。这种实践被广泛认为是安全的,即使出现少量标注错误,由于这些微小误差对模型最终性能的影响微不足道,且现有后门攻击需要攻击者具备污染训练图像的能力。然而,本文提出清洁图像后门攻击,揭示后门仍可通过修改少量错误标签(无需修改训练图像)实现注入。具体而言,在我们的攻击中,攻击者首先寻找一种触发器特征,将训练图像分为具有该特征与不具有该特征的两部分。随后,攻击者将前一部分图像的标签篡改为后门类别。当目标模型在污染数据上训练后,后门最终被植入模型。在推理阶段,攻击者可通过两种方式激活后门:轻微修改输入图像以获得触发器特征,或直接输入天然具有触发器特征的图像。我们通过大量实验验证了所提攻击的有效性和实用性。实验结果表明,我们的攻击严重威胁了图像分类模型的公平性与鲁棒性,且必须警惕外包标注中的错误标签。