The commercialization of diffusion models, renowned for their ability to generate high-quality images that are often indistinguishable from real ones, brings forth potential copyright concerns. Although attempts have been made to impede unauthorized access to copyrighted material during training and to subsequently prevent DMs from generating copyrighted images, the effectiveness of these solutions remains unverified. This study explores the vulnerabilities associated with copyright protection in DMs by introducing a backdoor data poisoning attack (SilentBadDiffusion) against text-to-image diffusion models. Our attack method operates without requiring access to or control over the diffusion model's training or fine-tuning processes; it merely involves the insertion of poisoning data into the clean training dataset. This data, comprising poisoning images equipped with prompts, is generated by leveraging the powerful capabilities of multimodal large language models and text-guided image inpainting techniques. Our experimental results and analysis confirm the method's effectiveness. By integrating a minor portion of non-copyright-infringing stealthy poisoning data into the clean dataset-rendering it free from suspicion-we can prompt the finetuned diffusion models to produce copyrighted content when activated by specific trigger prompts. These findings underline potential pitfalls in the prevailing copyright protection strategies and underscore the necessity for increased scrutiny and preventative measures against the misuse of DMs.
翻译:扩散模型因其生成与真实图像难以区分的高质量图像能力而闻名,其商业化应用带来了潜在的版权问题。尽管已有尝试在训练过程中阻止对受版权保护材料的未授权访问,并进而防止扩散模型生成受版权保护的图像,但这些解决方案的有效性尚未得到验证。本研究通过引入一种针对文本到图像扩散模型的后门数据投毒攻击(SilentBadDiffusion),探讨了扩散模型在版权保护方面的脆弱性。我们的攻击方法无需访问或控制扩散模型的训练或微调过程,只需将投毒数据插入干净的训练数据集。这些数据由配备提示词的投毒图像组成,借助多模态大语言模型和文本引导的图像修复技术的强大能力生成。实验结果表明,该方法有效。通过在干净数据集中整合少量不构成版权侵权的隐秘投毒数据(使其免于怀疑),我们能够促使微调后的扩散模型在特定触发提示词激活时生成受版权保护的内容。这些发现揭示了当前版权保护策略中潜在的问题,并强调了对扩散模型滥用行为加强审查和预防措施的必要性。