The transferability of adversarial examples can be exploited to launch black-box attacks. However, adversarial examples often present poor transferability. To alleviate this issue, by observing that the diversity of inputs can boost transferability, input regularization based methods are proposed, which craft adversarial examples by combining several transformed inputs. We reveal that input regularization based methods make resultant adversarial examples biased towards flat extreme regions. Inspired by this, we propose an attack called flatness-aware adversarial attack (FAA) which explicitly adds a flatness-aware regularization term in the optimization target to promote the resultant adversarial examples towards flat extreme regions. The flatness-aware regularization term involves gradients of samples around the resultant adversarial examples but optimizing gradients requires the evaluation of Hessian matrix in high-dimension spaces which generally is intractable. To address the problem, we derive an approximate solution to circumvent the construction of Hessian matrix, thereby making FAA practical and cheap. Extensive experiments show the transferability of adversarial examples crafted by FAA can be considerably boosted compared with state-of-the-art baselines.
翻译:对抗样本的可迁移性可用于发起黑盒攻击。然而,对抗样本往往呈现出较差的迁移性。为了缓解这一问题,通过观察到输入多样性可以提升迁移性,研究者提出了基于输入正则化的方法,这些方法通过组合多种变换后的输入来生成对抗样本。我们发现,基于输入正则化的方法使得生成的对抗样本偏向于平坦极值区域。受此启发,我们提出一种名为平坦性感知对抗攻击(FAA)的方法,该方法在优化目标中显式添加平坦性感知正则化项,以引导生成的对抗样本趋向平坦极值区域。该正则化项涉及对抗样本周围样本的梯度,但梯度优化需要在高维空间中计算海森矩阵,这通常难以实现。为解决此问题,我们推导出一种近似解来规避海森矩阵的构建,从而使FAA方法既实用又高效。大量实验表明,与当前最先进的基线方法相比,FAA生成的对抗样本的迁移性得到显著提升。