TIT-Score：基于文本-图像-文本一致性评估长提示驱动的文本-图像对齐 (TIT-Score: Evaluating Long-Prompt Based Text-to-Image Alignment via Text-to-Image-to-Text Consistency)

With the rapid advancement of large multimodal models (LMMs), recent text-to-image (T2I) models can generate high-quality images and demonstrate great alignment to short prompts. However, they still struggle to effectively understand and follow long and detailed prompts, displaying inconsistent generation. To address this challenge, we introduce LPG-Bench, a comprehensive benchmark for evaluating long-prompt-based text-to-image generation. LPG-Bench features 200 meticulously crafted prompts with an average length of over 250 words, approaching the input capacity of several leading commercial models. Using these prompts, we generate 2,600 images from 13 state-of-the-art models and further perform comprehensive human-ranked annotations. Based on LPG-Bench, we observe that state-of-the-art T2I alignment evaluation metrics exhibit poor consistency with human preferences on long-prompt-based image generation. To address the gap, we introduce a novel zero-shot metric based on text-to-image-to-text consistency, termed TIT, for evaluating long-prompt-generated images. The core concept of TIT is to quantify T2I alignment by directly comparing the consistency between the raw prompt and the LMM-produced description on the generated image, which includes an efficient score-based instantiation TIT-Score and a large-language-model (LLM) based instantiation TIT-Score-LLM. Extensive experiments demonstrate that our framework achieves superior alignment with human judgment compared to CLIP-score, LMM-score, etc., with TIT-Score-LLM attaining a 7.31% absolute improvement in pairwise accuracy over the strongest baseline. LPG-Bench and TIT methods together offer a deeper perspective to benchmark and foster the development of T2I models. All resources will be made publicly available.

翻译：随着大型多模态模型（LMMs）的快速发展，近期的文本到图像（T2I）模型已能生成高质量图像，并在短提示上展现出良好的对齐能力。然而，它们在有效理解和遵循冗长、详细的提示方面仍存在困难，表现出不一致的生成效果。为应对这一挑战，我们引入了LPG-Bench，一个用于评估基于长提示的文本到图像生成的综合性基准。LPG-Bench包含200个精心构建的提示，平均长度超过250词，接近多个主流商业模型的输入容量上限。利用这些提示，我们从13个前沿模型中生成了2,600张图像，并进一步进行了全面的人工排序标注。基于LPG-Bench，我们观察到当前最先进的T2I对齐评估指标在基于长提示的图像生成任务上与人类偏好的一致性较差。为弥补这一差距，我们引入了一种基于文本-图像-文本一致性的新型零样本度量方法，称为TIT，用于评估长提示生成的图像。TIT的核心思想是通过直接比较原始提示与LMM对生成图像所产生描述之间的一致性来量化T2I对齐度，具体包括一个高效的基于分数的实例化TIT-Score和一个基于大语言模型（LLM）的实例化TIT-Score-LLM。大量实验表明，相较于CLIP-score、LMM-score等方法，我们的框架与人类判断实现了更优的对齐，其中TIT-Score-LLM在成对准确率上相比最强基线获得了7.31%的绝对提升。LPG-Bench与TIT方法共同为基准测试和推动T2I模型的发展提供了更深入的视角。所有资源将公开提供。

相关内容

MoDELS

关注 45

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

【NeurIPS2021】用于文本图表示学习的 GNN 嵌套 Transformer 模型：GraphFormers

专知会员服务

46+阅读 · 2021年11月24日

UCM《机器学习导论笔记》，80页pdf CSE176 Introduction to Machine Learning

专知会员服务

31+阅读 · 2021年9月29日

Linux导论，Introduction to Linux，96页ppt

专知会员服务

82+阅读 · 2020年7月26日

FlowQA: Grasping Flow in History for Conversational Machine Comprehension

专知会员服务

34+阅读 · 2019年10月18日