Black-box adversarial attacks have shown strong potential to subvert machine learning models. Existing black-box adversarial attacks craft the adversarial examples by iteratively querying the target model and/or leveraging the transferability of a local surrogate model. Whether such attack can succeed remains unknown to the adversary when empirically designing the attack. In this paper, to our best knowledge, we take the first step to study a new paradigm of adversarial attacks -- certifiable black-box attack that can guarantee the attack success rate of the crafted adversarial examples. Specifically, we revise the randomized smoothing to establish novel theories for ensuring the attack success rate of the adversarial examples. To craft the adversarial examples with the certifiable attack success rate (CASR) guarantee, we design several novel techniques, including a randomized query method to query the target model, an initialization method with smoothed self-supervised perturbation to derive certifiable adversarial examples, and a geometric shifting method to reduce the perturbation size of the certifiable adversarial examples for better imperceptibility. We have comprehensively evaluated the performance of the certifiable black-box attack on CIFAR10 and ImageNet datasets against different levels of defenses. Both theoretical and experimental results have validated the effectiveness of the proposed certifiable attack.
翻译:黑盒对抗攻击在颠覆机器学习模型方面展现出巨大潜力。现有黑盒对抗攻击通过迭代查询目标模型和/或利用本地替代模型的迁移性来生成对抗样本。然而在经验性设计攻击时,攻击者无法预知此类攻击能否成功。本文首次提出一种新型对抗攻击范式——可认证黑盒攻击,该范式能够保证所生成对抗样本的攻击成功率。具体而言,我们改进了随机平滑理论,建立了确保对抗样本攻击成功率的全新理论框架。为生成具有可认证攻击成功率(Certifiable Attack Success Rate, CASR)保证的对抗样本,我们设计了多项创新技术:包括用于查询目标模型的随机查询方法、通过平滑自监督扰动初始化获得可认证对抗样本的方法,以及通过几何移位降低可认证对抗样本扰动幅度以提升隐蔽性的方法。我们在CIFAR10与ImageNet数据集上,针对不同防御强度全面评估了可认证黑盒攻击的性能。理论分析与实验结果均验证了所提可认证攻击的有效性。