While Vision-Language Models (VLMs) have shown strong visual reasoning capabilities, their spatial reasoning abilities remain largely constrained to the observed images and text-oriented chain-of-thought. They often struggle to infer unobserved layouts, maintain cross-view consistency, and reason from alternative viewpoints when only limited egocentric observations are available. In this work, we study this problem as thinking with imagination, where a VLM actively acquires imagined visual evidence by interacting with a world simulator during reasoning. We propose Astra, an agentic spatial reasoning framework that empowers VLMs with action-conditioned visual imagination. Specifically, Astra couples Astra-VL, an RL-trained VLM policy, with Astra-WM, a Bagel-based world simulator that generates novel-view observations from context images and natural-language camera motions. To provide reliable imagined evidence, Astra-WM is trained with view consistency tuning to improve pose and content consistency across views. In the RL stage, we propose a world-simulator-in-the-loop two-phase RL curriculum to stabilize tool-use exploration and advance the model's ability to invoke the simulator only when imagined observations improve over direct answering. Experiments demonstrate that both the world simulator and the agentic policy are necessary: Astra-WM improves simulator-augmented Gemini-3-Flash on MMSI-Bench from 45.1 to 49.5, while Astra-VL improves the Qwen3-VL backbone from 29.8 to 38.8 on MMSI-Bench and from 36.8 to 42.7 on MindCube. These results show that imagined observations can provide useful spatial evidence, but effective world-model-augmented reasoning requires learning when, where, and how to imagine.


翻译:尽管视觉-语言模型(Vision-Language Models, VLMs)已展现出强大的视觉推理能力,但其空间推理能力仍高度受限:仅能基于观测图像和文本导向的思维链进行推理。当仅能获取有限的第一人称观测时,它们往往难以推断未观测布局、维持跨视角一致性,以及从替代视角进行推理。本研究将此类问题定义为“思考与想象”(thinking with imagination),即VLM在推理过程中通过与世界模拟器交互,主动获取想象中的视觉证据。我们提出Astra——一种智能空间推理框架,赋予VLM基于动作条件的视觉想象能力。具体而言,Astra将基于强化学习(RL)训练的VLM策略Astra-VL与基于“贝果”(Bagel)的世界模拟器Astra-WM相结合,后者能从上下文图像和自然语言描述的相机运动生成新视角观测。为提供可靠的想象证据,Astra-WM通过视角一致性训练(view consistency tuning)提升跨视角的姿态与内容一致性。在强化学习阶段,我们提出一种“世界模拟器在环”的两阶段RL课程,以稳定工具使用探索并提升模型仅在想象观测优于直接回答时调用模拟器的能力。实验表明,世界模拟器与智能策略二者缺一不可:Astra-WM将模拟器增强的Gemini-3-Flash在MMSI-Bench上的性能从45.1提升至49.5;而Astra-VL将Qwen3-VL骨干模型在MMSI-Bench上的性能从29.8提升至38.8,在MindCube上从36.8提升至42.7。这些结果表明:想象观测能提供有效的空间证据,但有效的世界模型增强推理需要学习何时、何地以及如何想象。

0
下载
关闭预览

相关内容

从感知到推理:深度思考赋能多模态大语言模型
专知会员服务
26+阅读 · 2025年11月19日
【NeurIPS2023】大型语言模型是视觉推理协调器
专知会员服务
30+阅读 · 2023年10月24日
【博士论文】视觉语言交互中的视觉推理研究
专知会员服务
65+阅读 · 2021年12月1日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
Arxiv
0+阅读 · 6月15日
VIP会员
最新内容
俄乌无人机战争的六大启示
专知会员服务
8+阅读 · 8月3日
《无人机空中监控:通信实验洞察》
专知会员服务
6+阅读 · 8月3日
从采集到决策:美军视角下的战术情报范式重构
《履带式无人地面战车技术发展现状》
专知会员服务
6+阅读 · 8月2日
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
6+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
13+阅读 · 7月31日
相关资讯
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员