Text-to-image diffusion models suffer from the risk of generating outdated, copyrighted, incorrect, and biased content. While previous methods have mitigated the issues on a small scale, it is essential to handle them simultaneously in larger-scale real-world scenarios. We propose a two-stage method, Editing Massive Concepts In Diffusion Models (EMCID). The first stage performs memory optimization for each individual concept with dual self-distillation from text alignment loss and diffusion noise prediction loss. The second stage conducts massive concept editing with multi-layer, closed form model editing. We further propose a comprehensive benchmark, named ImageNet Concept Editing Benchmark (ICEB), for evaluating massive concept editing for T2I models with two subtasks, free-form prompts, massive concept categories, and extensive evaluation metrics. Extensive experiments conducted on our proposed benchmark and previous benchmarks demonstrate the superior scalability of EMCID for editing up to 1,000 concepts, providing a practical approach for fast adjustment and re-deployment of T2I diffusion models in real-world applications.
翻译:文本到图像扩散模型存在生成过时、受版权保护、错误及有偏见内容的风险。尽管先前的方法已在较小规模上缓解了这些问题,但在大规模现实场景中同时处理它们至关重要。我们提出了一种两阶段方法,即编辑扩散模型中的海量概念(EMCID)。第一阶段通过文本对齐损失和扩散噪声预测损失的双重自蒸馏对每个概念进行记忆优化。第二阶段采用多层、封闭形式的模型编辑进行海量概念编辑。我们进一步提出了一个综合基准,名为ImageNet概念编辑基准(ICEB),用于评估T2I模型的海量概念编辑,包含两个子任务、自由形式提示、海量概念类别以及广泛的评估指标。在我们提出的基准和先前基准上进行的大量实验表明,EMCID在编辑多达1000个概念方面具有卓越的可扩展性,为T2I扩散模型在实际应用中的快速调整和重新部署提供了实用方法。