Certified defense methods against adversarial perturbations have been recently investigated in the black-box setting with a zeroth-order (ZO) perspective. However, these methods suffer from high model variance with low performance on high-dimensional datasets due to the ineffective design of the denoiser and are limited in their utilization of ZO techniques. To this end, we propose a certified ZO preprocessing technique for removing adversarial perturbations from the attacked image in the black-box setting using only model queries. We propose a robust UNet denoiser (RDUNet) that ensures the robustness of black-box models trained on high-dimensional datasets. We propose a novel black-box denoised smoothing (DS) defense mechanism, ZO-RUDS, by prepending our RDUNet to the black-box model, ensuring black-box defense. We further propose ZO-AE-RUDS in which RDUNet followed by autoencoder (AE) is prepended to the black-box model. We perform extensive experiments on four classification datasets, CIFAR-10, CIFAR-10, Tiny Imagenet, STL-10, and the MNIST dataset for image reconstruction tasks. Our proposed defense methods ZO-RUDS and ZO-AE-RUDS beat SOTA with a huge margin of $35\%$ and $9\%$, for low dimensional (CIFAR-10) and with a margin of $20.61\%$ and $23.51\%$ for high-dimensional (STL-10) datasets, respectively.
翻译:针对对抗性扰动的认证防御方法近期在黑盒设置中从零阶(ZO)视角得到了研究。然而,由于去噪器设计的低效性,这些方法在高维数据集上存在模型方差大、性能低的问题,并且对ZO技术的利用有限。为此,我们提出一种认证的ZO预处理技术,仅通过模型查询即可在黑盒设置中从受攻击图像中移除对抗性扰动。我们提出一种鲁棒的UNet去噪器(RDUNet),确保训练于高维数据集上的黑盒模型的鲁棒性。通过将RDUNet前置到黑盒模型前,我们提出一种新颖的黑盒去噪平滑(DS)防御机制ZO-RUDS,实现黑盒防御。进一步,我们提出ZO-AE-RUDS方法,其中将RDUNet后接自编码器(AE)前置到黑盒模型前。我们在四个分类数据集(CIFAR-10、CIFAR-10、Tiny Imagenet、STL-10)以及用于图像重建任务的MNIST数据集上进行了大量实验。我们提出的防御方法ZO-RUDS和ZO-AE-RUDS在低维数据集(CIFAR-10)上分别以$35\%$和$9\%$的显著优势超越现有最优方法(SOTA),在高维数据集(STL-10)上分别以$20.61\%$和$23.51\%$的优势超越。