While Multimodal Large Language Models demonstrate impressive semantic capabilities, they often suffer from spatial blindness, struggling with fine-grained geometric reasoning and physical dynamics. Existing solutions typically rely on explicit 3D modalities or complex geometric scaffolding, which are limited by data scarcity and generalization challenges. In this work, we propose a paradigm shift by leveraging the implicit spatial prior within large-scale video generation models. We posit that to synthesize temporally coherent videos, these models inherently learn robust 3D structural priors and physical laws. We introduce VEGA-3D (Video Extracted Generative Awareness), a plug-and-play framework that repurposes a pre-trained video diffusion model as a Latent World Simulator. By extracting spatiotemporal features from intermediate noise levels and integrating them with semantic representations via a token-level adaptive gated fusion mechanism, we enrich MLLMs with dense geometric cues without explicit 3D supervision. Extensive experiments across 3D scene understanding, spatial reasoning, and embodied manipulation benchmarks demonstrate that our method outperforms state-of-the-art baselines, validating that generative priors provide a scalable foundation for physical-world understanding. Code is publicly available at https://github.com/H-EmbodVis/VEGA-3D.


翻译:尽管多模态大语言模型展现出令人瞩目的语义能力,但它们常受空间盲点所困,难以进行细粒度几何推理和物理动力学理解。现有解决方案通常依赖显式三维模态或复杂几何架构,受限于数据稀缺和泛化挑战。本研究提出范式转变,通过利用大规模视频生成模型中的隐式空间先验。我们提出假设:为合成时间连贯视频,这些模型内在地学习了鲁棒的三维结构先验和物理定律。我们引入VEGA-3D(视频提取的生成感知),一个即插即用框架,将预训练视频扩散模型重新用作潜在世界模拟器。通过从中间噪声水平提取时空特征,并借助令牌级自适应门控融合机制将其与语义表征整合,我们无需显式三维监督即可为MLLMs注入密集几何线索。在三维场景理解、空间推理和具身操作基准测试上的大量实验表明,我们的方法优于现有最优基线,验证了生成先验为物理世界理解提供了可扩展基础。代码已开源:https://github.com/H-EmbodVis/VEGA-3D。

0
下载
关闭预览

相关内容

从感知到推理:深度思考赋能多模态大语言模型
专知会员服务
26+阅读 · 2025年11月19日
绝对干货!NLP预训练模型:从transformer到albert
新智元
15+阅读 · 2019年11月10日
深度图像先验:无需学习即可生成新图像
论智
45+阅读 · 2017年12月4日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
23+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
13+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
反制无人机:乌克兰提供的五点启示
专知会员服务
1+阅读 · 今天14:52
《各指挥层级均亟需红队能力》报告
专知会员服务
2+阅读 · 今天14:46
《远程传感对战机架次生成影响的仿真研究》80页
《航电任务系统框架(FAMOS)》50页报告
专知会员服务
4+阅读 · 9月22日
《对抗行动中的人工智能与自主性》智库报告
专知会员服务
7+阅读 · 9月22日
《从数据到胜利:战争中的分析优势之争》
专知会员服务
9+阅读 · 9月22日
战争不仅需要机器人:人类仍不可或缺
专知会员服务
5+阅读 · 9月21日
《描绘美国防部创新基础设施的未来蓝图》100页
专知会员服务
10+阅读 · 9月21日
相关VIP内容
从感知到推理:深度思考赋能多模态大语言模型
专知会员服务
26+阅读 · 2025年11月19日
相关基金
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
23+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
13+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员