Self-supervised learning (SSL) is a prevalent approach for encoding data representations. Using a pre-trained SSL image encoder and subsequently training a downstream classifier, impressive performance can be achieved on various tasks with very little labeled data. The growing adoption of SSL has led to an increase in security research on SSL encoders and associated Trojan attacks. Trojan attacks embedded in SSL encoders can operate covertly, spreading across multiple users and devices. The presence of backdoor behavior in Trojaned encoders can inadvertently be inherited by downstream classifiers, making it even more difficult to detect and mitigate the threat. Although current Trojan detection methods in supervised learning can potentially safeguard SSL downstream classifiers, identifying and addressing triggers in the SSL encoder before its widespread dissemination is a challenging task. This challenge arises because downstream tasks might be unknown, dataset labels may be unavailable, and the original unlbeled training dataset might be inaccessible during Trojan detection in SSL encoders. We introduce SSL-Cleanse as a solution to identify and mitigate backdoor threats in SSL encoders. We evaluated SSL-Cleanse on various datasets using 1200 encoders, achieving an average detection success rate of 82.2% on ImageNet-100. After mitigating backdoors, on average, backdoored encoders achieve 0.3% attack success rate without great accuracy loss, proving the effectiveness of SSL-Cleanse.
翻译:自监督学习(SSL)是一种用于编码数据表示的常用方法。通过使用预训练的SSL图像编码器并训练下游分类器,可以在少量标注数据上实现不同任务的出色性能。随着SSL的广泛应用,针对SSL编码器的安全研究及相关木马攻击日益增多。嵌入在SSL编码器中的木马攻击可隐蔽运作,跨多个用户和设备传播。被植入木马的编码器中的后门行为可能被下游分类器无意继承,从而进一步增加了检测和缓解该威胁的难度。尽管现有的监督学习中的木马检测方法在一定程度上能保障SSL下游分类器的安全,但在SSL编码器广泛传播前识别和处理触发器仍是一项挑战。这一挑战源于:在SSL编码器的木马检测过程中,下游任务可能未知、数据集标签可能不可用、且原始无标签训练数据集可能无法获取。我们提出SSL-Cleanse作为识别和缓解SSL编码器中后门威胁的解决方案。我们使用1200个编码器在多个数据集上评估了SSL-Cleanse,在ImageNet-100上实现了平均82.2%的检测成功率。在缓解后门后,被植入后门的编码器平均攻击成功率降至0.3%且精度损失极小,验证了SSL-Cleanse的有效性。