Although promising, existing defenses against query-based attacks share a common limitation: they offer increased robustness against attacks at the price of a considerable accuracy drop on clean samples. In this work, we show how to efficiently establish, at test-time, a solid tradeoff between robustness and accuracy when mitigating query-based attacks. Given that these attacks necessarily explore low-confidence regions, our insight is that activating dedicated defenses, such as RND (Qin et al., NeuRIPS 2021) and Random Image Transformations (Xie et al., ICLR 2018), only for low-confidence inputs is sufficient to prevent them. Our approach is independent of training and supported by theory. We verify the effectiveness of our approach for various existing defenses by conducting extensive experiments on CIFAR-10, CIFAR-100, and ImageNet. Our results confirm that our proposal can indeed enhance these defenses by providing better tradeoffs between robustness and accuracy when compared to state-of-the-art approaches while being completely training-free.
翻译:尽管现有针对查询攻击的防御方法展现出前景,但它们普遍存在一个局限性:在提升对抗攻击鲁棒性的同时,会显著降低干净样本的准确率。本研究提出了一种在测试阶段高效建立鲁棒性与准确率之间稳定权衡的方法。鉴于此类攻击必然探索低置信度区域,我们的核心洞见是:仅对低置信度输入激活专用防御机制(如RND(Qin等,NeurIPS 2021)和随机图像变换(Xie等,ICLR 2018)),即可有效阻止攻击。该方法不依赖训练过程且具有理论支撑。通过在CIFAR-10、CIFAR-100和ImageNet上的大量实验,我们验证了该方法对多种现有防御的有效性。结果表明,与现有最优方法相比,我们的方案能够在完全无需训练的情况下,为这些防御提供更优的鲁棒性与准确率权衡。