PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models

Text-to-image (T2I) models have made substantial progress in generating images from textual prompts. However, they frequently fail to produce images consistent with physical commonsense, a vital capability for applications in world simulation and everyday tasks. Current T2I evaluation benchmarks focus on metrics such as accuracy, bias, and safety, neglecting the evaluation of models' internal knowledge, particularly physical commonsense. To address this issue, we introduce PhyBench, a comprehensive T2I evaluation dataset comprising 700 prompts across 4 primary categories: mechanics, optics, thermodynamics, and material properties, encompassing 31 distinct physical scenarios. We assess 6 prominent T2I models, including proprietary models DALLE3 and Gemini, and demonstrate that incorporating physical principles into prompts enhances the models' ability to generate physically accurate images. Our findings reveal that: (1) even advanced models frequently err in various physical scenarios, except for optics; (2) GPT-4o, with item-specific scoring instructions, effectively evaluates the models' understanding of physical commonsense, closely aligning with human assessments; and (3) current T2I models are primarily focused on text-to-image translation, lacking profound reasoning regarding physical commonsense. We advocate for increased attention to the inherent knowledge within T2I models, beyond their utility as mere image generation tools. The code and data are available at https://github.com/OpenGVLab/PhyBench.

翻译：文本到图像（T2I）模型在根据文本提示生成图像方面取得了实质性进展。然而，它们经常无法生成与物理常识一致的图像，而这是世界模拟和日常任务应用中的一项关键能力。当前的T2I评估基准侧重于准确性、偏见和安全性等指标，忽视了对模型内部知识，特别是物理常识的评估。为解决这一问题，我们引入了PhyBench，一个全面的T2I评估数据集，包含700个提示，涵盖力学、光学、热力学和材料属性4个主要类别，涉及31个不同的物理场景。我们评估了6个主流的T2I模型，包括专有模型DALLE3和Gemini，并证明将物理原理融入提示可以增强模型生成物理准确图像的能力。我们的研究结果表明：（1）即使是先进模型也经常在各种物理场景中出错，光学场景除外；（2）配备特定项目评分指令的GPT-4o，能有效评估模型对物理常识的理解，其结果与人类评估高度一致；（3）当前的T2I模型主要侧重于文本到图像的转换，缺乏对物理常识的深度推理。我们主张应更多地关注T2I模型的内在知识，而不仅仅是将其视为图像生成工具。代码和数据可在 https://github.com/OpenGVLab/PhyBench 获取。

相关内容

MoDELS

关注 45

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

【CVPR 2022】一个完全无监督的框架，从噪声和部分测量中学习图像，Robust Equivariant Imaging: a fully unsupervised framework for learning to image

专知会员服务

25+阅读 · 2022年3月3日

【亚马逊-WWW2020】不解析,生成!用于面向任务的语义分析的序列到序列体系结构，Don't Parse, Generate! A Sequence to Sequence Architecture for Task-Oriented Semantic Parsing

专知会员服务

15+阅读 · 2020年2月1日

FlowQA: Grasping Flow in History for Conversational Machine Comprehension

专知会员服务

34+阅读 · 2019年10月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日