Real-world adversarial physical patches were shown to be successful in compromising state-of-the-art models in a variety of computer vision applications. Existing defenses that are based on either input gradient or features analysis have been compromised by recent GAN-based attacks that generate naturalistic patches. In this paper, we propose Jedi, a new defense against adversarial patches that is resilient to realistic patch attacks. Jedi tackles the patch localization problem from an information theory perspective; leverages two new ideas: (1) it improves the identification of potential patch regions using entropy analysis: we show that the entropy of adversarial patches is high, even in naturalistic patches; and (2) it improves the localization of adversarial patches, using an autoencoder that is able to complete patch regions from high entropy kernels. Jedi achieves high-precision adversarial patch localization, which we show is critical to successfully repair the images. Since Jedi relies on an input entropy analysis, it is model-agnostic, and can be applied on pre-trained off-the-shelf models without changes to the training or inference of the protected models. Jedi detects on average 90% of adversarial patches across different benchmarks and recovers up to 94% of successful patch attacks (Compared to 75% and 65% for LGS and Jujutsu, respectively).
翻译:现实世界中的对抗性物理补丁已被证明能够成功攻破多种计算机视觉应用中的先进模型。现有基于输入梯度或特征分析的防御方法,已被近期能够生成自然化补丁的GAN类攻击所突破。本文提出Jedi——一种针对对抗性补丁的新型防御方法,其对真实补丁攻击具有鲁棒性。Jedi从信息论视角解决补丁定位问题,并融合两项创新:(1)利用熵分析改进潜在补丁区域的识别:我们证明了对抗性补丁(包括自然化补丁)具有高熵特性;(2)通过自编码器实现高熵核区域的补丁完成,提升对抗性补丁的定位精度。Jedi实现了高精度的对抗性补丁定位,这对后续成功修复图像至关重要。由于Jedi依赖于输入熵分析,其具有模型无关性,可直接应用于预训练的现成模型,无需修改受保护模型的训练或推理过程。Jedi在不同基准测试中平均检测90%的对抗性补丁,并成功恢复高达94%的补丁攻击(相比之下,LGS与Jujutsu的恢复率分别为75%和65%)。