Since adversarial examples appeared and showed the catastrophic degradation they brought to DNN, many adversarial defense methods have been devised, among which adversarial training is considered the most effective. However, a recent work showed the inequality phenomena in $l_{\infty}$-adversarial training and revealed that the $l_{\infty}$-adversarially trained model is vulnerable when a few important pixels are perturbed by i.i.d. noise or occluded. In this paper, we propose a simple yet effective method called Input Gradient Distillation (IGD) to release the inequality phenomena in $l_{\infty}$-adversarial training. Experiments show that while preserving the model's adversarial robustness, compared to PGDAT, IGD decreases the $l_{\infty}$-adversarially trained model's error rate to inductive noise and inductive occlusion by up to 60\% and 16.53\%, and to noisy images in Imagenet-C by up to 21.11\%. Moreover, we formally explain why the equality of the model's saliency map can improve such robustness.
翻译:自对抗样本出现并展示其对深度神经网络造成的灾难性性能退化以来,研究者已设计出多种对抗防御方法,其中对抗训练被认为最为有效。然而,近期研究揭示了$l_{\infty}$-对抗训练中存在的不平等现象,并指出经$l_{\infty}$-对抗训练的模型在少量重要像素被独立同分布噪声扰动或遮挡时表现脆弱。本文提出一种简洁而高效的方法——输入梯度蒸馏(IGD),用于缓解$l_{\infty}$-对抗训练中的不平等现象。实验表明,在保持模型对抗鲁棒性的同时,相比PGDAT方法,IGD将$l_{\infty}$-对抗训练模型对归纳噪声与归纳遮挡的误差率分别降低高达60%和16.53%,对ImageNet-C噪声图像的误差率降低高达21.11%。此外,我们从理论层面解释了模型显著性图的平等性为何能提升此类鲁棒性。