Unrestricted adversarial attacks typically manipulate the semantic content of an image (e.g., color or texture) to create adversarial examples that are both effective and photorealistic, demonstrating their ability to deceive human perception and deep neural networks with stealth and success. However, current works usually sacrifice unrestricted degrees and subjectively select some image content to guarantee the photorealism of unrestricted adversarial examples, which limits its attack performance. To ensure the photorealism of adversarial examples and boost attack performance, we propose a novel unrestricted attack framework called Content-based Unrestricted Adversarial Attack. By leveraging a low-dimensional manifold that represents natural images, we map the images onto the manifold and optimize them along its adversarial direction. Therefore, within this framework, we implement Adversarial Content Attack based on Stable Diffusion and can generate high transferable unrestricted adversarial examples with various adversarial contents. Extensive experimentation and visualization demonstrate the efficacy of ACA, particularly in surpassing state-of-the-art attacks by an average of 13.3-50.4% and 16.8-48.0% in normally trained models and defense methods, respectively.
翻译:不受限对抗攻击通常通过操纵图像的语义内容(如颜色或纹理)来生成既有效又具有照片真实感的对抗样本,从而展示其欺骗人类感知和深度神经网络的能力,具有隐蔽性和成功性。然而,当前工作通常牺牲不受限程度,主观选择部分图像内容以保证不受限对抗样本的照片真实感,这限制了其攻击性能。为确保对抗样本的照片真实感并提升攻击性能,我们提出了一种新颖的不受限攻击框架,称为基于内容的不受限对抗攻击。通过利用表示自然图像的低维流形,我们将图像映射到流形上,并沿其对抗方向进行优化。因此,在此框架内,我们基于Stable Diffusion实现了对抗内容攻击,并能够生成具有多种对抗内容的高迁移性不受限对抗样本。大量实验和可视化证明了ACA的有效性,特别是在正常训练模型和防御方法中,它平均分别超越了现有最先进攻击13.3-50.4%和16.8-48.0%。