Vision-language models (VLMs) and generative world models are opening new opportunities for embodied navigation. VLMs are increasingly used as direct planners or trajectory predictors, while world models support look-ahead reasoning by imagining future views. Yet predicting a reliable trajectory from a single egocentric observation remains challenging. Current VLMs often generate unstable trajectories, and world models, though able to synthesize plausible futures, do not directly provide the grounded signals needed for navigation learning. This raises a central question: how can generated futures be turned into supervision for grounded trajectory prediction? We present WorldMAP, a teacher--student framework that converts world-model-generated futures into persistent semantic-spatial structure and planning-derived supervision. Its world-model-driven teacher builds semantic-spatial memory from generated videos, grounds task-relevant targets and obstacles, and produces trajectory pseudo-labels through explicit planning. A lightweight student with a multi-hypothesis trajectory head is then trained to predict navigation trajectories directly from vision-language inputs. On Target-Bench, WorldMAP achieves the best ADE and FDE among compared methods, reducing ADE by 18.0% and FDE by 42.1% relative to the best competing baseline, while lifting a small open-source VLM to DTW performance competitive with proprietary models. More broadly, the results suggest that, in embodied navigation, the value of world models may lie less in supplying action-ready imagined evidence than in synthesizing structured supervision for navigation learning.
翻译:视觉-语言模型(VLM)与生成式世界模型为正具身导航开辟了新机遇。VLM越来越多地被用作直接规划器或轨迹预测器,而世界模型则通过想象未来视图来支持前瞻性推理。然而,从单一自我中心观测中预测可靠轨迹仍具挑战。现有VLM常生成不稳定轨迹,而世界模型虽能合成合理未来状态,却无法直接提供导航学习所需的锚定信号。这引出一个核心问题:如何将生成的未来场景转化为有监督信号,用于基于锚定的轨迹预测?我们提出WorldMAP——一种教师-学生框架,可将世界模型生成的未来场景转化为持久语义-空间结构与规划导引监督。其世界模型驱动的教师模块从生成视频中构建语义-空间记忆,锚定任务相关目标与障碍物,并通过显式规划生成轨迹伪标签。随后训练具有多假设轨迹头的轻量学生模型,使其能直接从视觉-语言输入预测导航轨迹。在Target-Bench基准上,WorldMAP在对比方法中实现了最佳平均位移误差(ADE)与最终位移误差(FDE),相对最优基线分别降低ADE 18.0%、FDE 42.1%,同时使小型开源VLM的DTW性能达到与专有模型竞争的水平。更广泛而言,结果表明:在正具身导航中,世界模型的价值或许不在于提供可直接用于行动的想象证据,而在于合成为导航学习提供结构化监督。