The unprecedented photorealistic results achieved by recent text-to-image generative systems and their increasing use as plug-and-play content creation solutions make it crucial to understand their potential biases. In this work, we introduce three indicators to evaluate the realism, diversity and prompt-generation consistency of text-to-image generative systems when prompted to generate objects from across the world. Our indicators complement qualitative analysis of the broader impact of such systems by enabling automatic and efficient benchmarking of geographic disparities, an important step towards building responsible visual content creation systems. We use our proposed indicators to analyze potential geographic biases in state-of-the-art visual content creation systems and find that: (1) models have less realism and diversity of generations when prompting for Africa and West Asia than Europe, (2) prompting with geographic information comes at a cost to prompt-consistency and diversity of generated images, and (3) models exhibit more region-level disparities for some objects than others. Perhaps most interestingly, our indicators suggest that progress in image generation quality has come at the cost of real-world geographic representation. Our comprehensive evaluation constitutes a crucial step towards ensuring a positive experience of visual content creation for everyone.
翻译:近年来文本到图像生成系统取得了前所未有的逼真效果,且其作为即插即用的内容创作解决方案的应用日益广泛,这使得理解其潜在偏见变得至关重要。本研究提出三项指标,用于评估文本到图像生成系统在生成全球各地物体时的逼真度、多样性和提示词-生成内容一致性。通过实现地理差异的自动化高效基准测试,我们的指标补充了对这类系统更广泛影响的定性分析,这是构建负责任的视觉内容创作系统的重要步骤。我们运用所提出的指标分析当前最优视觉内容创作系统中的潜在地理偏见,发现:(1)当提示词涉及非洲和西亚地区时,模型生成的逼真度和多样性低于欧洲;(2)在提示词中加入地理信息会降低生成图像的提示一致性及多样性;(3)不同物体类别在模型中的地域差异程度不同。最有趣的是,我们的指标表明图像生成质量的进步是以牺牲真实世界的地理表征为代价的。这一全面评估为确保所有人获得良好的视觉内容创作体验迈出了关键一步。