Vision-language models (VLMs) and generative world models are opening new opportunities for embodied navigation. VLMs are increasingly used as direct planners or trajectory predictors, while world models support look-ahead reasoning by imagining future views. Yet predicting a reliable trajectory from a single egocentric observation remains challenging. Current VLMs often generate unstable trajectories, and world models, though able to synthesize plausible futures, do not directly provide the grounded signals needed for navigation learning. This raises a central question: how can generated futures be turned into supervision for grounded trajectory prediction? We present WorldMAP, a teacher--student framework that converts world-model-generated futures into persistent semantic-spatial structure and planning-derived supervision. Its world-model-driven teacher builds semantic-spatial memory from generated videos, grounds task-relevant targets and obstacles, and produces trajectory pseudo-labels through explicit planning. A lightweight student with a multi-hypothesis trajectory head is then trained to predict navigation trajectories directly from vision-language inputs. On Target-Bench, WorldMAP achieves the best ADE and FDE among compared methods, reducing ADE by 18.0% and FDE by 42.1% relative to the best competing baseline, while lifting a small open-source VLM to DTW performance competitive with proprietary models. More broadly, the results suggest that, in embodied navigation, the value of world models may lie less in supplying action-ready imagined evidence than in synthesizing structured supervision for navigation learning.


翻译:视觉-语言模型(VLM)与生成式世界模型为正具身导航开辟了新机遇。VLM越来越多地被用作直接规划器或轨迹预测器,而世界模型则通过想象未来视图来支持前瞻性推理。然而,从单一自我中心观测中预测可靠轨迹仍具挑战。现有VLM常生成不稳定轨迹,而世界模型虽能合成合理未来状态,却无法直接提供导航学习所需的锚定信号。这引出一个核心问题:如何将生成的未来场景转化为有监督信号,用于基于锚定的轨迹预测?我们提出WorldMAP——一种教师-学生框架,可将世界模型生成的未来场景转化为持久语义-空间结构与规划导引监督。其世界模型驱动的教师模块从生成视频中构建语义-空间记忆,锚定任务相关目标与障碍物,并通过显式规划生成轨迹伪标签。随后训练具有多假设轨迹头的轻量学生模型,使其能直接从视觉-语言输入预测导航轨迹。在Target-Bench基准上,WorldMAP在对比方法中实现了最佳平均位移误差(ADE)与最终位移误差(FDE),相对最优基线分别降低ADE 18.0%、FDE 42.1%,同时使小型开源VLM的DTW性能达到与专有模型竞争的水平。更广泛而言,结果表明:在正具身导航中,世界模型的价值或许不在于提供可直接用于行动的想象证据,而在于合成为导航学习提供结构化监督。

0
下载
关闭预览

相关内容

从看见到认知世界:视觉世界模型综述
专知会员服务
17+阅读 · 5月17日
视觉语言建模遇见遥感:模型、数据集与前景展望
专知会员服务
17+阅读 · 2025年5月21日
自动驾驶的世界模型综述
专知会员服务
47+阅读 · 2025年1月22日
《面向视觉语言地理基础模型》综述
专知会员服务
47+阅读 · 2024年6月15日
探索视觉语言模型的前沿:当前方法和未来方向的综述
专知会员服务
49+阅读 · 2024年4月12日
自然语言处理中的语言模型预训练方法
PaperWeekly
14+阅读 · 2018年10月21日
视觉里程计:特征点法之全面梳理
计算机视觉life
12+阅读 · 2017年8月2日
视觉里程计:起源、优势、对比、应用
计算机视觉life
18+阅读 · 2017年7月17日
国家自然科学基金
8+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
VIP会员
相关主题
最新内容
对抗环境下超视距目标打击的情报支援
专知会员服务
8+阅读 · 7月22日
《无人机对海面作战影响评估》
专知会员服务
15+阅读 · 7月21日
印度精确打击与指挥架构的断层
专知会员服务
7+阅读 · 7月20日
美空军AI完成F-16战斗机自主空战历史性试飞
专知会员服务
8+阅读 · 7月20日
相关基金
国家自然科学基金
8+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员