Adversarial examples (AE) with good transferability enable practical black-box attacks on diverse target models, where insider knowledge about the target models is not required. Previous methods often generate AE with no or very limited transferability; that is, they easily overfit to the particular architecture and feature representation of the source, white-box model and the generated AE barely work for target, black-box models. In this paper, we propose a novel approach to enhance AE transferability using Gradient Norm Penalty (GNP). It drives the loss function optimization procedure to converge to a flat region of local optima in the loss landscape. By attacking 11 state-of-the-art (SOTA) deep learning models and 6 advanced defense methods, we empirically show that GNP is very effective in generating AE with high transferability. We also demonstrate that it is very flexible in that it can be easily integrated with other gradient based methods for stronger transfer-based attacks.
翻译:对抗样本(AE)若具有良好的可迁移性,可在无需了解目标模型内部知识的情况下实现实际的黑盒攻击。以往方法生成的对抗样本通常可迁移性极低甚至为零,即容易过拟合源白盒模型的特定架构与特征表示,导致生成的对抗样本几乎无法攻击目标黑盒模型。本文提出一种利用梯度范数罚项(GNP)增强对抗样本可迁移性的新方法。该方法驱动损失函数优化过程收敛至损失景观中局部最优解的平坦区域。通过攻击11种最先进的深度学习模型和6种先进防御方法,实验表明GNP在生成具有高可迁移性的对抗样本方面非常有效。我们还证明该方法具有高度灵活性,可轻松与其他基于梯度的攻击方法集成以增强基于迁移的攻击效果。