AI-Generated Content (AIGC) is gaining great popularity, with many emerging commercial services and applications. These services leverage advanced generative models, such as latent diffusion models and large language models, to generate creative content (e.g., realistic images and fluent sentences) for users. The usage of such generated content needs to be highly regulated, as the service providers need to ensure the users do not violate the usage policies (e.g., abuse for commercialization, generating and distributing unsafe content). A promising solution to achieve this goal is watermarking, which adds unique and imperceptible watermarks on the content for service verification and attribution. Numerous watermarking approaches have been proposed recently. However, in this paper, we show that an adversary can easily break these watermarking mechanisms. Specifically, we consider two possible attacks. (1) Watermark removal: the adversary can easily erase the embedded watermark from the generated content and then use it freely bypassing the regulation of the service provider. (2) Watermark forging: the adversary can create illegal content with forged watermarks from another user, causing the service provider to make wrong attributions. We propose Warfare, a unified methodology to achieve both attacks in a holistic way. The key idea is to leverage a pre-trained diffusion model for content processing and a generative adversarial network for watermark removal or forging. We evaluate Warfare on different datasets and embedding setups. The results prove that it can achieve high success rates while maintaining the quality of the generated content. Compared to existing diffusion model-based attacks, Warfare is 5,050~11,000x faster.
翻译:人工智能生成内容(AIGC)正日益普及,涌现出众多新兴的商业服务与应用。这些服务利用先进的生成模型(如潜在扩散模型和大语言模型)为用户生成创意内容(例如逼真的图像和流畅的句子)。此类生成内容的使用需要严格监管,因为服务提供者需确保用户不违反使用政策(例如滥用商业化、生成和传播不安全内容)。实现这一目标的一种有前景的方案是水印技术,即向内容中添加独特且不可感知的水印,用于服务验证和归属溯源。近期已提出大量水印方法。然而,本文表明,攻击者可轻易破坏这些水印机制。具体而言,我们考虑两种可能的攻击:(1)水印去除:攻击者可轻松擦除生成内容中嵌入的水印,进而自由使用该内容以绕过服务提供者的监管。(2)水印伪造:攻击者可利用伪造自其他用户的水印创建非法内容,导致服务提供者做出错误的归属判定。我们提出统一方法论Warfare,以整体方式实现上述两种攻击。其核心思想是借助预训练的扩散模型进行内容处理,并利用生成对抗网络实现水印去除或伪造。我们在不同数据集和嵌入设置下评估了Warfare。结果表明,该方法能在保持生成内容质量的同时实现高成功率。与现有基于扩散模型的攻击相比,Warfare的速度提升5,050至11,000倍。