Despite the recent advances in personalized text-to-image (P-T2I) generative models, subject-driven T2I remains challenging. The primary bottlenecks include 1) Intensive training resource requirements, 2) Hyper-parameter sensitivity leading to inconsistent outputs, and 3) Balancing the intricacies of novel visual concept and composition alignment. We start by re-iterating the core philosophy of T2I diffusion models to address the above limitations. Predominantly, contemporary subject-driven T2I approaches hinge on Latent Diffusion Models (LDMs), which facilitate T2I mapping through cross-attention layers. While LDMs offer distinct advantages, P-T2I methods' reliance on the latent space of these diffusion models significantly escalates resource demands, leading to inconsistent results and necessitating numerous iterations for a single desired image. Recently, ECLIPSE has demonstrated a more resource-efficient pathway for training UnCLIP-based T2I models, circumventing the need for diffusion text-to-image priors. Building on this, we introduce $\lambda$-ECLIPSE. Our method illustrates that effective P-T2I does not necessarily depend on the latent space of diffusion models. $\lambda$-ECLIPSE achieves single, multi-subject, and edge-guided T2I personalization with just 34M parameters and is trained on a mere 74 GPU hours using 1.6M image-text interleaved data. Through extensive experiments, we also establish that $\lambda$-ECLIPSE surpasses existing baselines in composition alignment while preserving concept alignment performance, even with significantly lower resource utilization.
翻译:尽管近来个性化文生图(P-T2I)生成模型取得了进展,但主体驱动的文生图仍然具有挑战性。主要瓶颈包括:1)高强度的训练资源需求;2)超参数敏感性导致输出不一致;3)难以平衡新颖视觉概念与构图对齐的复杂性。为克服上述局限,我们首先重新阐述了T2I扩散模型的核心原理。当前主流的主体驱动T2I方法依赖于潜在扩散模型(LDM),通过交叉注意力层实现T2I映射。虽然LDM具有明显优势,但P-T2I方法对其扩散模型潜在空间的依赖显著增加了资源需求,导致结果不一致,且生成单张目标图像需多次迭代。近期,ECLIPSE展示了训练基于UnCLIP的T2I模型更具资源效率的路径,避免了扩散文本到图像先验的需求。在此基础上,我们提出λ-ECLIPSE。该方法表明,有效的P-T2I并不必然依赖扩散模型的潜在空间。λ-ECLIPSE仅需3400万参数,在74个GPU小时的训练时长内,使用160万图文交织数据即可实现单主体、多主体及边缘引导的T2I个性化。大量实验证明,即便在资源利用率显著降低的情况下,λ-ECLIPSE在保持概念对齐性能的同时,在构图对齐方面仍超越现有基线模型。