Denoising diffusion models have shown remarkable potential in various generation tasks. The open-source large-scale text-to-image model, Stable Diffusion, becomes prevalent as it can generate realistic artistic or facial images with personalization through fine-tuning on a limited number of new samples. However, this has raised privacy concerns as adversaries can acquire facial images online and fine-tune text-to-image models for malicious editing, leading to baseless scandals, defamation, and disruption to victims' lives. Prior research efforts have focused on deriving adversarial loss from conventional training processes for facial privacy protection through adversarial perturbations. However, existing algorithms face two issues: 1) they neglect the image-text fusion module, which is the vital module of text-to-image diffusion models, and 2) their defensive performance is unstable against different attacker prompts. In this paper, we propose the Adversarial Decoupling Augmentation Framework (ADAF), addressing these issues by targeting the image-text fusion module to enhance the defensive performance of facial privacy protection algorithms. ADAF introduces multi-level text-related augmentations for defense stability against various attacker prompts. Concretely, considering the vision, text, and common unit space, we propose Vision-Adversarial Loss, Prompt-Robust Augmentation, and Attention-Decoupling Loss. Extensive experiments on CelebA-HQ and VGGFace2 demonstrate ADAF's promising performance, surpassing existing algorithms.
翻译:去噪扩散模型在各种生成任务中展现出显著潜力。开源的大规模文本到图像模型Stable Diffusion因其能够通过对有限数量的新样本进行微调,生成逼真的艺术或面部图像而流行。然而,这也引发了隐私担忧:攻击者可以获取在线面部图像,并微调文本到图像模型以进行恶意编辑,导致无端丑闻、诽谤,并扰乱受害者生活。先前的研究工作主要集中在通过对抗性扰动从传统训练过程中衍生对抗性损失来保护面部隐私。然而,现有算法面临两个问题:1)它们忽视了文本到图像扩散模型的关键模块——图像文本融合模块;2)它们对不同攻击者提示的防御性能不稳定。本文提出了对抗解耦增强框架(ADAF),通过针对图像文本融合模块来解决这些问题,从而增强面部隐私保护算法的防御性能。ADAF引入了多级文本相关增强,以针对各种攻击者提示实现防御稳定性。具体而言,考虑到视觉、文本和公共单元空间,我们提出了视觉对抗损失、提示鲁棒增强和注意力解耦损失。在CelebA-HQ和VGGFace2上的广泛实验表明,ADAF表现优越,超越了现有算法。