Diffusion models arise as a powerful generative tool recently. Despite the great progress, existing diffusion models mainly focus on uni-modal control, i.e., the diffusion process is driven by only one modality of condition. To further unleash the users' creativity, it is desirable for the model to be controllable by multiple modalities simultaneously, e.g., generating and editing faces by describing the age (text-driven) while drawing the face shape (mask-driven). In this work, we present Collaborative Diffusion, where pre-trained uni-modal diffusion models collaborate to achieve multi-modal face generation and editing without re-training. Our key insight is that diffusion models driven by different modalities are inherently complementary regarding the latent denoising steps, where bilateral connections can be established upon. Specifically, we propose dynamic diffuser, a meta-network that adaptively hallucinates multi-modal denoising steps by predicting the spatial-temporal influence functions for each pre-trained uni-modal model. Collaborative Diffusion not only collaborates generation capabilities from uni-modal diffusion models, but also integrates multiple uni-modal manipulations to perform multi-modal editing. Extensive qualitative and quantitative experiments demonstrate the superiority of our framework in both image quality and condition consistency.
翻译:扩散模型近年来作为一种强大的生成工具崭露头角。尽管取得了巨大进展,现有的扩散模型主要聚焦于单模态控制,即扩散过程仅由单一模态的条件驱动。为进一步释放用户的创造力,模型应能同时受多种模态控制,例如通过描述年龄(文本驱动)生成并编辑人脸,同时绘制脸型(掩码驱动)。本文提出协同扩散(Collaborative Diffusion),其中预训练的单模态扩散模型通过协作实现多模态人脸生成与编辑,而无需重新训练。我们的关键洞察在于:不同模态驱动的扩散模型在潜在去噪步骤上具有内在互补性,可由此建立双向连接。具体而言,我们提出动态扩散器(dynamic diffuser),一种元网络,通过预测每个预训练单模态模型的时空影响函数,自适应地生成多模态去噪步骤。协同扩散不仅整合了单模态扩散模型的生成能力,还融合了多种单模态操作以实现多模态编辑。大量定性与定量实验表明,我们的框架在图像质量和条件一致性上均具有优越性。