Personalized text-to-image generation using diffusion models has recently been proposed and attracted lots of attention. Given a handful of images containing a novel concept (e.g., a unique toy), we aim to tune the generative model to capture fine visual details of the novel concept and generate photorealistic images following a text condition. We present a plug-in method, named ViCo, for fast and lightweight personalized generation. Specifically, we propose an image attention module to condition the diffusion process on the patch-wise visual semantics. We introduce an attention-based object mask that comes almost at no cost from the attention module. In addition, we design a simple regularization based on the intrinsic properties of text-image attention maps to alleviate the common overfitting degradation. Unlike many existing models, our method does not finetune any parameters of the original diffusion model. This allows more flexible and transferable model deployment. With only light parameter training (~6% of the diffusion U-Net), our method achieves comparable or even better performance than all state-of-the-art models both qualitatively and quantitatively.
翻译:基于扩散模型的个性化文本到图像生成近期被提出并受到广泛关注。给定少量包含新概念(如独特玩具)的图像,我们旨在调整生成模型以捕捉该新概念的精细视觉细节,并依据文本条件生成逼真图像。我们提出一种名为ViCo的即插即用方法,用于快速轻量的个性化生成。具体而言,我们设计了一个图像注意力模块,通过分块视觉语义条件约束扩散过程;引入一个来自注意力模块且几乎零成本的基于注意力的对象掩码;此外,我们基于文本-图像注意力图的内在属性设计了一种简单正则化方法,以缓解常见的过拟合退化问题。与现有诸多模型不同,本方法无需微调原始扩散模型的任何参数,这使得模型部署更加灵活且可迁移。仅通过少量参数训练(约占扩散U-Net的6%),我们的方法在定性与定量性能上均能达到甚至超越所有最先进模型。