Deep neural networks have proven to be vulnerable to adversarial attacks in the form of adding specific perturbations on images to make wrong outputs. Designing stronger adversarial attack methods can help more reliably evaluate the robustness of DNN models. To release the harbor burden and improve the attack performance, auto machine learning (AutoML) has recently emerged as one successful technique to help automatically find the near-optimal adversarial attack strategy. However, existing works about AutoML for adversarial attacks only focus on $L_{\infty}$-norm-based perturbations. In fact, semantic perturbations attract increasing attention due to their naturalnesses and physical realizability. To bridge the gap between AutoML and semantic adversarial attacks, we propose a novel method called multi-objective evolutionary search of variable-length composite semantic perturbations (MES-VCSP). Specifically, we construct the mathematical model of variable-length composite semantic perturbations, which provides five gradient-based semantic attack methods. The same type of perturbation in an attack sequence is allowed to be performed multiple times. Besides, we introduce the multi-objective evolutionary search consisting of NSGA-II and neighborhood search to find near-optimal variable-length attack sequences. Experimental results on CIFAR10 and ImageNet datasets show that compared with existing methods, MES-VCSP can obtain adversarial examples with a higher attack success rate, more naturalness, and less time cost.
翻译:深度神经网络已被证明易受对抗攻击的影响,即通过在图像上添加特定扰动来产生错误输出。设计更强的对抗攻击方法有助于更可靠地评估深度神经网络模型的鲁棒性。为减轻人工负担并提升攻击性能,自动机器学习(AutoML)最近作为一种成功的技术出现,可帮助自动寻找接近最优的对抗攻击策略。然而,现有关于AutoML对抗攻击的研究仅关注基于$L_{\infty}$范数的扰动。实际上,语义扰动因其自然性和物理可实现性而日益受到关注。为弥合AutoML与语义对抗攻击之间的差距,我们提出一种新方法,称为变长复合语义扰动的多目标进化搜索(MES-VCSP)。具体而言,我们构建了变长复合语义扰动的数学模型,该模型提供五种基于梯度的语义攻击方法。攻击序列中允许同一类型的扰动多次执行。此外,我们引入由NSGA-II和邻域搜索组成的多目标进化搜索,以寻找接近最优的变长攻击序列。在CIFAR10和ImageNet数据集上的实验结果表明,与现有方法相比,MES-VCSP能够获得攻击成功率更高、自然性更好且时间成本更低的对抗样本。