Recent advances in large text-conditional image generative models such as Stable Diffusion, Midjourney, and DALL-E 3 have revolutionized the field of image generation, allowing users to produce high-quality, realistic images from textual prompts. While these developments have enhanced artistic creation and visual communication, they also present an underexplored attack opportunity: the possibility of inducing biases by an adversary into the generated images for malicious intentions, e.g., to influence society and spread propaganda. In this paper, we demonstrate the possibility of such a bias injection threat by an adversary who backdoors such models with a small number of malicious data samples; the implemented backdoor is activated when special triggers exist in the input prompt of the backdoored models. On the other hand, the model's utility is preserved in the absence of the triggers, making the attack highly undetectable. We present a novel framework that enables efficient generation of poisoning samples with composite (multi-word) triggers for such an attack. Our extensive experiments using over 1 million generated images and against hundreds of fine-tuned models demonstrate the feasibility of the presented backdoor attack. We illustrate how these biases can bypass conventional detection mechanisms, highlighting the challenges in proving the existence of biases within operational constraints. Our cost analysis confirms the low financial barrier to executing such attacks, underscoring the need for robust defensive strategies against such vulnerabilities in text-to-image generation models.
翻译:近期,以Stable Diffusion、Midjourney和DALL-E 3为代表的大型文本条件图像生成模型取得了突破性进展,彻底改变了图像生成领域,使用户能够通过文本提示生成高质量、逼真的图像。尽管这些进展促进了艺术创作与视觉交流,但也带来了尚未被充分探索的攻击可能性:攻击者可能出于恶意目的(例如影响社会舆论、传播宣传内容)在生成图像中诱导偏见。本文通过实验证明,攻击者仅需少量恶意数据样本即可对此类模型植入后门,当被植入后门的模型输入提示中存在特定触发器时,后门将被激活。另一方面,在无触发器存在时模型功能保持正常,使得该攻击极具隐蔽性。我们提出了一种新颖的框架,能够为此类攻击高效生成包含复合(多词)触发器的投毒样本。基于超过100万张生成图像、针对数百个微调模型开展的广泛实验,验证了所提后门攻击的可行性。我们展示了此类偏见如何规避常规检测机制,凸显了在操作约束下证明偏见存在的挑战。成本分析证实了实施此类攻击的经济门槛较低,这强调了对文本到图像生成模型中此类漏洞建立强健防御策略的迫切需求。