Recent text-to-image diffusion models are able to learn and synthesize images containing novel, personalized concepts (e.g., their own pets or specific items) with just a few examples for training. This paper tackles two interconnected issues within this realm of personalizing text-to-image diffusion models. First, current personalization techniques fail to reliably extend to multiple concepts -- we hypothesize this to be due to the mismatch between complex scenes and simple text descriptions in the pre-training dataset (e.g., LAION). Second, given an image containing multiple personalized concepts, there lacks a holistic metric that evaluates performance on not just the degree of resemblance of personalized concepts, but also whether all concepts are present in the image and whether the image accurately reflects the overall text description. To address these issues, we introduce Gen4Gen, a semi-automated dataset creation pipeline utilizing generative models to combine personalized concepts into complex compositions along with text-descriptions. Using this, we create a dataset called MyCanvas, that can be used to benchmark the task of multi-concept personalization. In addition, we design a comprehensive metric comprising two scores (CP-CLIP and TI-CLIP) for better quantifying the performance of multi-concept, personalized text-to-image diffusion methods. We provide a simple baseline built on top of Custom Diffusion with empirical prompting strategies for future researchers to evaluate on MyCanvas. We show that by improving data quality and prompting strategies, we can significantly increase multi-concept personalized image generation quality, without requiring any modifications to model architecture or training algorithms.
翻译:近期文本到图像扩散模型仅需少量训练样例便能习得并合成包含新颖个性化概念(如用户的宠物或特定物品)的图像。本文针对文本到图像扩散模型个性化领域中的两个相互关联的问题展开研究:其一,当前的个性化技术难以可靠地扩展至多概念场景——我们推测这是由于预训练数据集(如LAION)中复杂场景与简单文本描述之间的失配所致;其二,给定包含多个个性化概念的图像,目前缺乏一种能全面评估性能的整体指标——该指标不仅需度量个性化概念的相似度,还应评估图像中是否包含所有概念以及图像是否准确反映整体文本描述。为解决这些问题,我们提出Gen4Gen——一种半自动化数据集创建流水线,利用生成模型将个性化概念融入复杂组合场景并生成对应的文本描述。基于此流水线,我们创建了MyCanvas数据集,可用于多概念个性化任务的基准测试。此外,我们设计了一套包含两个分数(CP-CLIP和TI-CLIP)的综合评估指标,以更精确地量化多概念个性化文本到图像扩散方法的性能。我们基于Custom Diffusion构建了一个简单基线模型,并搭配实证性提示策略,供未来研究者在MyCanvas上进行评估。实验表明,通过提升数据质量与优化提示策略,我们能够在无需修改模型架构或训练算法的前提下,显著提高多概念个性化图像生成的综合质量。