Reversible adversarial examples (RAE) combine adversarial attacks and reversible data-hiding technology on a single image to prevent illegal access. Most RAE studies focus on achieving white-box attacks. In this paper, we propose a novel framework to generate reversible adversarial examples, which combines a novel beam search based black-box attack and reversible data hiding with grayscale invariance (RDH-GI). This RAE uses beam search to evaluate the adversarial gain of historical perturbations and guide adversarial perturbations. After the adversarial examples are generated, the framework RDH-GI embeds the secret data that can be recovered losslessly. Experimental results show that our method can achieve an average Peak Signal-to-Noise Ratio (PSNR) of at least 40dB compared to source images with limited query budgets. Our method can also achieve a targeted black-box reversible adversarial attack for the first time.
翻译:可逆对抗样本(RAE)将对抗攻击与可逆数据隐藏技术结合于单张图像,以防止非法访问。现有RAE研究主要聚焦于白盒攻击。本文提出一种新颖的可逆对抗样本生成框架,该框架融合了基于束搜索的黑盒攻击与灰度不变性可逆数据隐藏(RDH-GI)。该RAE通过束搜索评估历史扰动的对抗增益并指导对抗扰动生成。在对抗样本生成后,RDH-GI框架嵌入可无损恢复的隐秘数据。实验结果表明,在有限查询预算下,本方法相较原始图像的平均峰值信噪比(PSNR)可达至少40dB。同时,本方法首次实现了目标导向的黑盒可逆对抗攻击。