NewtPhys: Do Foundation Models Understand Newtonian Physics?

Previous work has evaluated physics reasoning in foundation models using synthetic or semi-synthetic scenes and visual question-answering tasks. However, these benchmarks emphasize high-level events and lack the visual fidelity required to assess true low-level Newtonian understanding. We introduce NewtPhys, a 4D physically annotated dataset built from multiview images of real-world scenes with physics-grounded simulations. The dataset provides dense, fine-grained annotations across timesteps -- including 3D forces and amodal per-pixel quantities covering physics, tracking, semantics and geometry -- bridging the gap between simplistic synthetic setups and realistic visual complexity. Using NewtPhys, we systematically evaluate 56 VLMs, including 54 open-weight models and 2 closed-source frontier models, and 10 VFMs and reveal limitations in low-level physics reasoning. Beyond benchmarking, our dataset enables future research in physics-grounded vision and the development of next-generation physics-aware evaluations. Code and datasets are available at https://astra-vision.github.io/NewtPhys.

翻译：先前研究通过合成或半合成场景及视觉问答任务，评估了基础模型在物理推理方面的能力。然而，这些基准测试侧重于高层事件，缺乏评估真正低层牛顿理解所需的视觉保真度。我们提出NewtPhys，这是一个基于真实世界场景多视图图像和物理驱动模拟构建的四维物理标注数据集。该数据集在时间步上提供密集的细粒度标注——包括三维力与非模态逐像素量（涵盖物理、跟踪、语义和几何）——弥合了简单合成环境与真实视觉复杂度之间的鸿沟。利用NewtPhys，我们系统评估了56个视觉语言模型（含54个开源模型与2个闭源前沿模型）及10个视觉基础模型，揭示了它们在低层物理推理中的局限性。除基准测试外，本数据集还可支撑未来基于物理的视觉研究，以及下一代物理感知评估方法的开发。代码和数据集详见https://astra-vision.github.io/NewtPhys。

相关内容

MoDELS

关注 45

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

【牛津博士论文】无监督物体学习（Unsupervised Object Learning）

专知会员服务

14+阅读 · 2025年11月30日

【斯坦福博士论文】多模态基础模型：从科学理解到科学发现

专知会员服务

31+阅读 · 2025年11月9日

【博士论文】弥合多模态基础模型与世界模型之间的鸿沟

专知会员服务

33+阅读 · 2025年10月9日

【新书】基于物理的模拟

专知会员服务

23+阅读 · 2025年7月25日