Adversarial example is a rising way of protecting facial privacy security from deepfake modification. To prevent massive facial images from being illegally modified by various deepfake models, it is essential to design a universal deepfake disruptor. However, existing works treat deepfake disruption as an End-to-End process, ignoring the functional difference between feature extraction and image reconstruction, which makes it difficult to generate a cross-model universal disruptor. In this work, we propose a novel Feature-Output ensemble UNiversal Disruptor (FOUND) against deepfake networks, which explores a new opinion that considers attacking feature extractors as the more critical and general task in deepfake disruption. We conduct an effective two-stage disruption process. We first disrupt multi-model feature extractors through multi-feature aggregation and individual-feature maintenance, and then develop a gradient-ensemble algorithm to enhance the disruption effect by simplifying the complex optimization problem of disrupting multiple End-to-End models. Extensive experiments demonstrate that FOUND can significantly boost the disruption effect against ensemble deepfake benchmark models. Besides, our method can fast obtain a cross-attribute, cross-image, and cross-model universal deepfake disruptor with only a few training images, surpassing state-of-the-art universal disruptors in both success rate and efficiency.
翻译:对抗样本是一种新兴的面部隐私安全保护方法,可有效防御深度伪造篡改。为防止大量面部图像被各类深度伪造模型非法修改,设计通用型深度伪造干扰器至关重要。然而,现有方法将深度伪造干扰视为端到端过程,忽略特征提取与图像重建的功能差异,导致难以生成跨模型的通用干扰器。本文提出一种面向深度伪造网络的**特征-输出集成通用干扰器**(FOUND),该机制探索了新的观点,即认为攻击特征提取器是深度伪造干扰中更为关键且普适的任务。我们通过有效的两阶段干扰流程实现:首先通过多特征聚合与单特征维护机制破坏多模型特征提取器,进而开发梯度集成算法,通过简化多端到端模型的复杂优化问题来增强干扰效果。大量实验表明,FOUND可显著提升对集成深度伪造基准模型的干扰效果。此外,本方法仅需少量训练图像即可快速获得具有跨属性、跨图像、跨模型特性的通用深度伪造干扰器,在成功率和效率上均超越现有最优通用干扰器。