With the widespread application of deep learning across various domains, concerns about its security have grown significantly. Among these, backdoor attacks pose a serious security threat to deep neural networks (DNNs). In recent years, backdoor attacks on neural networks have become increasingly sophisticated, aiming to compromise the security and trustworthiness of models by implanting hidden, unauthorized functionalities or triggers, leading to misleading predictions or behaviors. To make triggers less perceptible and imperceptible, various invisible backdoor attacks have been proposed. However, most of them only consider invisibility in the spatial domain, making it easy for recent defense methods to detect the generated toxic images.To address these challenges, this paper proposes an invisible backdoor attack called DEBA. DEBA leverages the mathematical properties of Singular Value Decomposition (SVD) to embed imperceptible backdoors into models during the training phase, thereby causing them to exhibit predefined malicious behavior under specific trigger conditions. Specifically, we first perform SVD on images, and then replace the minor features of trigger images with those of clean images, using them as triggers to ensure the effectiveness of the attack. As minor features are scattered throughout the entire image, the major features of clean images are preserved, making poisoned images visually indistinguishable from clean ones. Extensive experimental evaluations demonstrate that DEBA is highly effective, maintaining high perceptual quality and a high attack success rate for poisoned images. Furthermore, we assess the performance of DEBA under existing defense measures, showing that it is robust and capable of significantly evading and resisting the effects of these defense measures.
翻译:随着深度学习在各个领域的广泛应用,其安全性问题日益受到关注。其中,后门攻击对深度神经网络(DNN)构成了严重的安全威胁。近年来,针对神经网络的后门攻击技术愈发复杂,旨在通过植入隐藏的、未经授权的功能或触发器来破坏模型的可靠性与可信度,从而导致模型产生误导性预测或行为。为使触发器更不易察觉,研究者提出了多种隐形后门攻击方法。然而,现有方法大多仅考虑空间域中的不可见性,这使得近期防御手段易于检测出生成的毒性图像。为应对这些挑战,本文提出了一种名为DEBA的隐形后门攻击方法。DEBA利用奇异值分解(SVD)的数学特性,在训练阶段向模型中嵌入难以感知的后门,从而使其在特定触发条件下表现出预设的恶意行为。具体而言,我们首先对图像进行SVD分解,然后用干净图像的次要特征替换触发图像的次要特征,将其作为触发器以确保攻击的有效性。由于次要特征散布于整幅图像中,而干净图像的主要特征得以保留,因此被污染的图像在视觉上与干净图像难以区分。大量实验评估表明,DEBA具有很高的有效性,能够使被污染图像保持较高的感知质量并实现高攻击成功率。此外,我们在现有防御措施下评估了DEBA的性能,结果表明该攻击具有较强的鲁棒性,能够显著规避并抵抗这些防御手段的影响。