Recent advancements in text-to-image diffusion models have yielded impressive results in generating realistic and diverse images. However, these models still struggle with complex prompts, such as those that involve numeracy and spatial reasoning. This work proposes to enhance prompt understanding capabilities in diffusion models. Our method leverages a pretrained large language model (LLM) for grounded generation in a novel two-stage process. In the first stage, the LLM generates a scene layout that comprises captioned bounding boxes from a given prompt describing the desired image. In the second stage, a novel controller guides an off-the-shelf diffusion model for layout-grounded image generation. Both stages utilize existing pretrained models without additional model parameter optimization. Our method significantly outperforms the base diffusion model and several strong baselines in accurately generating images according to prompts that require various capabilities, doubling the generation accuracy across four tasks on average. Furthermore, our method enables instruction-based multi-round scene specification and can handle prompts in languages not supported by the underlying diffusion model. We anticipate that our method will unleash users' creativity by accurately following more complex prompts.
翻译:摘要:近期文本到图像扩散模型的进展在生成逼真且多样化的图像方面取得了显著成果。然而,这些模型在处理复杂提示(例如涉及计数和空间推理的提示)时仍存在困难。本文提出一种增强扩散模型提示理解能力的方法。该方法通过一种新颖的两阶段流程,利用预训练的大型语言模型(LLM)进行接地生成。在第一阶段,LLM根据描述期望图像的给定提示生成包含带标题边界框的场景布局。在第二阶段,一种新型控制器引导现成的扩散模型进行布局接地图像生成。两个阶段均使用现有预训练模型,无需额外的模型参数优化。与基础扩散模型及数个强基线方法相比,我们的方法在根据需要多种能力的提示准确生成图像方面显著更优,平均在四项任务上将生成准确率提升了一倍。此外,该方法支持基于指令的多轮场景指定,并能处理底层扩散模型未覆盖的语言提示。我们预计,该方法将通过精准遵循更复杂的提示,充分释放用户的创造力。