Recent research has highlighted the vulnerability of Deep Neural Networks (DNNs) against data poisoning attacks. These attacks aim to inject poisoning samples into the models' training dataset such that the trained models have inference failures. While previous studies have executed different types of attacks, one major challenge that greatly limits their effectiveness is the uncertainty of the re-training process after the injection of poisoning samples, including the re-training initialization or algorithms. To address this challenge, we propose a novel attack method called ''Sharpness-Aware Data Poisoning Attack (SAPA)''. In particular, it leverages the concept of DNNs' loss landscape sharpness to optimize the poisoning effect on the worst re-trained model. It helps enhance the preservation of the poisoning effect, regardless of the specific retraining procedure employed. Extensive experiments demonstrate that SAPA offers a general and principled strategy that significantly enhances various types of poisoning attacks.
翻译:近期研究揭示了深度神经网络(DNNs)面对数据投毒攻击的脆弱性。此类攻击旨在向模型训练数据集中注入投毒样本,导致训练后的模型产生推理错误。尽管先前研究已实现多种攻击类型,但极大制约其有效性的一个主要挑战在于,投毒样本注入后的重新训练过程存在不确定性,包括重新训练初始化或算法的不确定性。为应对这一挑战,我们提出了一种名为"锐度感知数据投毒攻击(SAPA)"的新型攻击方法。该方法利用深度神经网络损失景观锐度(loss landscape sharpness)的概念,优化投注效果在最劣重新训练模型上的表现。这有助于增强投毒效应的持久性,无论后续采用何种具体的重新训练流程。大量实验表明,SAPA提供了一种通用且原则化的策略,显著增强了多种类型的投毒攻击。