As a dominant force in text-to-image generation tasks, Diffusion Probabilistic Models (DPMs) face a critical challenge in controllability, struggling to adhere strictly to complex, multi-faceted instructions. In this work, we aim to address this alignment challenge for conditional generation tasks. First, we provide an alternative view of state-of-the-art DPMs as a way of inverting advanced Vision-Language Models (VLMs). With this formulation, we naturally propose a training-free approach that bypasses the conventional sampling process associated with DPMs. By directly optimizing images with the supervision of discriminative VLMs, the proposed method can potentially achieve a better text-image alignment. As proof of concept, we demonstrate the pipeline with the pre-trained BLIP-2 model and identify several key designs for improved image generation. To further enhance the image fidelity, a Score Distillation Sampling module of Stable Diffusion is incorporated. By carefully balancing the two components during optimization, our method can produce high-quality images with near state-of-the-art performance on T2I-Compbench.
翻译:作为文本到图像生成任务中的主导力量,扩散概率模型(DPMs)在可控性方面面临关键挑战,难以严格遵循复杂且多方面的指令。在这项工作中,我们旨在解决条件生成任务中的对齐问题。首先,我们提供了一种替代视角,将最先进的DPMs视为逆向高级视觉-语言模型(VLMs)的方法。基于这一公式,我们自然地提出了一种无需训练的方法,绕过了与DPMs相关的传统采样过程。通过直接优化图像并借助判别式VLMs的监督,所提出的方法有望实现更好的文本-图像对齐。作为概念验证,我们使用预训练的BLIP-2模型展示了该流程,并确定了若干关键设计以改进图像生成。为了进一步增强图像保真度,我们整合了稳定扩散的分数蒸馏采样模块。通过在优化过程中仔细平衡这两个组件,我们的方法能够生成高质量图像,并在T2I-Compbench上达到近乎最先进的性能。