Large text-to-image models have shown remarkable performance in synthesizing high-quality images. In particular, the subject-driven model makes it possible to personalize the image synthesis for a specific subject, e.g., a human face or an artistic style, by fine-tuning the generic text-to-image model with a few images from that subject. Nevertheless, misuse of subject-driven image synthesis may violate the authority of subject owners. For example, malicious users may use subject-driven synthesis to mimic specific artistic styles or to create fake facial images without authorization. To protect subject owners against such misuse, recent attempts have commonly relied on adversarial examples to indiscriminately disrupt subject-driven image synthesis. However, this essentially prevents any benign use of subject-driven synthesis based on protected images. In this paper, we take a different angle and aim at protection without sacrificing the utility of protected images for general synthesis purposes. Specifically, we propose GenWatermark, a novel watermark system based on jointly learning a watermark generator and a detector. In particular, to help the watermark survive the subject-driven synthesis, we incorporate the synthesis process in learning GenWatermark by fine-tuning the detector with synthesized images for a specific subject. This operation is shown to largely improve the watermark detection accuracy and also ensure the uniqueness of the watermark for each individual subject. Extensive experiments validate the effectiveness of GenWatermark, especially in practical scenarios with unknown models and text prompts (74% Acc.), as well as partial data watermarking (80% Acc. for 1/4 watermarking). We also demonstrate the robustness of GenWatermark to two potential countermeasures that substantially degrade the synthesis quality.
翻译:大规模文本到图像模型在合成高质量图像方面展现出显著性能。特别是主体驱动模型,通过使用特定主体的少量图像微调通用文本到图像模型,能针对特定主体(如人脸或艺术风格)实现个性化图像合成。然而,主体驱动图像合成的滥用可能侵犯主体所有者的权利。例如,恶意用户可能利用主体驱动合成模仿特定艺术风格,或在未经授权的情况下生成虚假人脸图像。为保护主体所有者免受此类滥用,近期尝试普遍依赖对抗性样本无差别地破坏主体驱动图像合成。但这本质上阻碍了基于受保护图像进行主体驱动合成的任何良性使用。本文从不同角度出发,旨在实现保护的同时不牺牲受保护图像用于通用合成目的的可用性。具体而言,我们提出GenWatermark——一种基于联合学习水印生成器与检测器的新型水印系统。为使水印在主体驱动合成中留存,我们在GenWatermark学习中融入合成过程,通过使用特定主体的合成图像微调检测器。实验表明,该操作能大幅提升水印检测精度,并确保每个独立主体水印的唯一性。大量实验验证了GenWatermark的有效性,尤其在未知模型与文本提示的实际场景中(准确率74%),以及部分数据水印场景下(1/4数据水印准确率达80%)。我们还证明了GenWatermark对两种显著降低合成质量的潜在对抗措施的鲁棒性。