Although recent personalization methods have democratized high-resolution image synthesis by enabling swift concept acquisition with minimal examples and lightweight computation, they also present an exploitable avenue for high accessible backdoor attacks. This paper investigates a critical and unexplored aspect of text-to-image (T2I) diffusion models - their potential vulnerability to backdoor attacks via personalization. Our study focuses on a zero-day backdoor vulnerability prevalent in two families of personalization methods, epitomized by Textual Inversion and DreamBooth.Compared to traditional backdoor attacks, our proposed method can facilitate more precise, efficient, and easily accessible attacks with a lower barrier to entry. We provide a comprehensive review of personalization in T2I diffusion models, highlighting the operation and exploitation potential of this backdoor vulnerability. To be specific, by studying the prompt processing of Textual Inversion and DreamBooth, we have devised dedicated backdoor attacks according to the different ways of dealing with unseen tokens and analyzed the influence of triggers and concept images on the attack effect. Our empirical study has shown that the nouveau-token backdoor attack has better attack performance while legacy-token backdoor attack is potentially harder to defend.
翻译:尽管近期个性化方法通过少量样本与轻量计算实现了快速概念获取,极大推动了高分辨率图像合成的民主化进程,但这也为高度可及的后门攻击提供了可乘之机。本文研究了文本到图像扩散模型中一个关键且尚未探索的领域——通过个性化方法遭受后门攻击的潜在脆弱性。我们聚焦于以Textual Inversion和DreamBooth为代表的两种个性化方法中普遍存在的零日后门漏洞。与传统后门攻击相比,本研究所提出的方法能够以更低的准入门槛实现更精准、高效且易实施的攻击。我们系统综述了文本到图像扩散模型的个性化技术,重点揭示了该后门漏洞的运行机制与利用潜力。具体而言,通过分析Textual Inversion与DreamBooth的提示词处理流程,我们根据两者处理未见令牌的不同方式分别设计了针对性后门攻击,并系统考察了触发器与概念图像对攻击效果的影响。实证研究表明,新型令牌后门攻击具有更优的攻击性能,而传统令牌后门攻击则潜在地更难防御。