Recent advances in text-to-image diffusion models have achieved remarkable success in generating high-quality, realistic images from given text prompts. However, previous methods fail to perform accurate modality alignment between text concepts and generated images due to the lack of fine-level semantic guidance that successfully diagnoses the modality discrepancy. In this paper, we propose FineRewards to improve the alignment between text and images in text-to-image diffusion models by introducing two new fine-grained semantic rewards: the caption reward and the Semantic Segment Anything (SAM) reward. From the global semantic view, the caption reward generates a corresponding detailed caption that depicts all important contents in the synthetic image via a BLIP-2 model and then calculates the reward score by measuring the similarity between the generated caption and the given prompt. From the local semantic view, the SAM reward segments the generated images into local parts with category labels, and scores the segmented parts by measuring the likelihood of each category appearing in the prompted scene via a large language model, i.e., Vicuna-7B. Additionally, we adopt an assemble reward-ranked learning strategy to enable the integration of multiple reward functions to jointly guide the model training. Adapting results of text-to-image models on the MS-COCO benchmark show that the proposed semantic reward outperforms other baseline reward functions with a considerable margin on both visual quality and semantic similarity with the input prompt. Moreover, by adopting the assemble reward-ranked learning strategy, we further demonstrate that model performance is further improved when adapting under the unifying of the proposed semantic reward with the current image rewards.
翻译:近期,文本到图像扩散模型在根据给定文本提示生成高质量、逼真图像方面取得了显著成功。然而,现有方法由于缺乏能有效诊断模态差异的细粒度语义引导,无法准确实现文本概念与生成图像之间的模态对齐。为此,本文提出FineRewards,通过引入两种新的细粒度语义奖励——描述奖励和语义分割任意物体(SAM)奖励,来改善文本到图像扩散模型中文本与图像的对齐。从全局语义视角,描述奖励利用BLIP-2模型为合成图像生成包含所有重要内容的详细描述,进而通过计算生成描述与给定提示的相似度得出奖励分数。从局部语义视角,SAM奖励将生成图像分割为带有类别标签的局部区域,并利用大语言模型(即Vicuna-7B)测量每个类别在提示场景中出现似然度,对各区域进行评分。此外,我们采用聚合奖励排序学习策略,实现多种奖励函数的集成以联合指导模型训练。在MS-COCO基准上的文本到图像模型适配结果表明,所提出的语义奖励在视觉质量和与输入提示的语义相似度上均显著优于其他基线奖励函数。进一步地,通过采用聚合奖励排序学习策略,我们证明将所提语义奖励与现有图像奖励统一适配时,模型性能可得到进一步提升。