The conditional text-to-image diffusion models have garnered significant attention in recent years. However, the precision of these models is often compromised mainly for two reasons, ambiguous condition input and inadequate condition guidance over single denoising loss. To address the challenges, we introduce two innovative solutions. Firstly, we propose a Spatial Guidance Injector (SGI) which enhances conditional detail by encoding text inputs with precise annotation information. This method directly tackles the issue of ambiguous control inputs by providing clear, annotated guidance to the model. Secondly, to overcome the issue of limited conditional supervision, we introduce Diffusion Consistency Loss (DCL), which applies supervision on the denoised latent code at any given time step. This encourages consistency between the latent code at each time step and the input signal, thereby enhancing the robustness and accuracy of the output. The combination of SGI and DCL results in our Effective Controllable Network (ECNet), which offers a more accurate controllable end-to-end text-to-image generation framework with a more precise conditioning input and stronger controllable supervision. We validate our approach through extensive experiments on generation under various conditions, such as human body skeletons, facial landmarks, and sketches of general objects. The results consistently demonstrate that our method significantly enhances the controllability and robustness of the generated images, outperforming existing state-of-the-art controllable text-to-image models.
翻译:条件文生图扩散模型近年来引起了广泛关注。然而,这些模型的精度往往因两个主要原因而受到损害:模糊的条件输入和基于单一去噪损失的不足条件指导。为应对这些挑战,我们引入了两项创新解决方案。首先,我们提出空间引导注入器(SGI),通过将文本输入与精确的注释信息编码来增强条件细节。该方法通过向模型提供清晰、注释化的引导,直接解决了模糊控制输入问题。其次,为克服有限条件监督的局限,我们引入扩散一致性损失(DCL),该损失对任意时间步长的去噪潜变量代码施加监督。这促进了每个时间步长潜变量代码与输入信号之间的一致性,从而增强了输出的鲁棒性和准确性。SGI与DCL的结合产生了我们的有效可控网络(ECNet),它通过更精确的条件输入和更强的可控监督,提供了一种更准确、端到端的可控文本到图像生成框架。我们通过在人体骨骼、面部关键点和通用物体草图等多种条件下的生成实验验证了我们的方法。实验结果一致表明,我们的方法显著增强了生成图像的可控性和鲁棒性,性能优于现有的最先进可控文本到图像模型。