While numerous defense methods have been proposed to prohibit potential poisoning attacks from untrusted data sources, most research works only defend against specific attacks, which leaves many avenues for an adversary to exploit. In this work, we propose an efficient and robust training approach to defend against data poisoning attacks based on influence functions, named Healthy Influential-Noise based Training. Using influence functions, we craft healthy noise that helps to harden the classification model against poisoning attacks without significantly affecting the generalization ability on test data. In addition, our method can perform effectively when only a subset of the training data is modified, instead of the current method of adding noise to all examples that has been used in several previous works. We conduct comprehensive evaluations over two image datasets with state-of-the-art poisoning attacks under different realistic attack scenarios. Our empirical results show that HINT can efficiently protect deep learning models against the effect of both untargeted and targeted poisoning attacks.
翻译:摘要:尽管已有众多防御方法被提出以防范来自不可信数据源的潜在投毒攻击,但大多数研究工作仅能抵御特定攻击,这为攻击者留下了许多可乘之机。本文提出一种基于影响力函数的高效鲁棒训练方法,命名为基于健康影响力噪声的训练(Healthy Influential-Noise based Training)。通过影响力函数,我们生成有助于增强分类模型对抗投毒攻击能力的健康噪声,且对测试数据的泛化能力影响甚微。此外,不同于先前若干工作对全部样本添加噪声的现有方法,本方法仅在训练数据子集被修改时仍能有效运作。我们针对两种图像数据集,结合最先进的投毒攻击方法,在多类现实攻击场景下进行了全面评估。实验结果表明,HINT能有效保护深度学习模型免受非定向及定向投毒攻击的影响。