Recently, DALL-E, a multimodal transformer language model, and its variants, including diffusion models, have shown high-quality text-to-image generation capabilities. However, despite the realistic image generation results, there has not been a detailed analysis of how to evaluate such models. In this work, we investigate the visual reasoning capabilities and social biases of different text-to-image models, covering both multimodal transformer language models and diffusion models. First, we measure three visual reasoning skills: object recognition, object counting, and spatial relation understanding. For this, we propose PaintSkills, a compositional diagnostic evaluation dataset that measures these skills. Despite the high-fidelity image generation capability, a large gap exists between the performance of recent models and the upper bound accuracy in object counting and spatial relation understanding skills. Second, we assess the gender and skin tone biases by measuring the gender/skin tone distribution of generated images across various professions and attributes. We demonstrate that recent text-to-image generation models learn specific biases about gender and skin tone from web image-text pairs. We hope our work will help guide future progress in improving text-to-image generation models on visual reasoning skills and learning socially unbiased representations. Code and data: https://github.com/j-min/DallEval
翻译:近年来,DALL-E这一多模态Transformer语言模型及其变体(包括扩散模型)展示了高质量文本到图像生成能力。然而,尽管能生成逼真的图像,目前仍缺乏对这些模型评估方法的详细分析。本研究系统探究了不同文本到图像模型(涵盖多模态Transformer语言模型与扩散模型)的视觉推理能力与社会偏见。首先,我们测量了三种视觉推理技能:物体识别、物体计数和空间关系理解。为此,我们提出了PaintSkills,一个用于评估这些技能的组成性诊断数据集。尽管模型具备高保真图像生成能力,但在物体计数和空间关系理解技能上,现有模型性能与上界准确率之间存在显著差距。其次,我们通过分析生成图像中不同职业及属性的性别/肤色分布,评估了性别与肤色偏见。实验表明,当前文本到图像生成模型从网络图文对中学习了特定的性别与肤色偏见。我们期望这项工作能指导未来在提升文本到图像生成模型的视觉推理能力与学习无社会偏见表征方面的进展。代码与数据:https://github.com/j-min/DallEval