Vision-Language-Action models (VLAs) leverage large-scale vision-language pretraining for semantic robot control, but often lack explicit foresight into how robot actions change the scene. World-Action Models (WAMs) address this limitation by conditioning policies on predicted futures, yet existing approaches typically rely on computationally expensive video generation with substantial pixel-level redundancy. We present LaWAM, a Latent World Action Model that exposes predictive dynamics to robot policies through compact latent visual subgoals instead of reconstructed future video. At the core of LaWAM is a latent-action-conditioned Latent World Model (LaWM). We obtain LaWM by training a latent action model in the latent space of a pretrained vision foundation model and repurposing its forward decoder to predict future observation features for scene evolution. LaWAM then conditions action generation on these predicted latent visual subgoals to enable dynamics-aware robot control. LaWAM achieves state-of-the-art or competitive success rates (SRs) across LIBERO (98.6% SR), RoboTwin (91.22% SR), and real-world manipulation tasks while retaining low-latency inference. LaWAM runs in 187 ms per action-chunk prediction and achieves up to 24x lower wall-clock latency than pixel-space WAMs.


翻译:视觉-语言-行动模型利用大规模视觉-语言预训练实现语义级机器人控制,但通常缺乏对机器人动作如何改变场景的显式预判。世界行动模型通过将策略建立在预测的未来状态之上来解决这一局限,然而现有方法通常依赖计算代价高昂且存在大量像素级冗余的视频生成。我们提出LaWAM——一种潜在世界行动模型,该模型通过紧凑的潜在视觉子目标而非重建的未来视频,向机器人策略暴露预测性动力学信息。LaWAM的核心是一个受潜在动作条件约束的潜在世界模型。我们通过在预训练视觉基础模型的潜在空间中训练潜在动作模型,并重新利用其前向解码器预测用于场景演化的未来观测特征,从而获得LaWM。LaWAM进而将这些预测的潜在视觉子目标作为动作生成的条件,实现动力学感知的机器人控制。在LIBERO(98.6%成功率)、RoboTwin(91.22%成功率)和真实世界操作任务中,LaWAM达到了最优或具有竞争力的成功率,同时保持低延迟推理。LaWAM每次动作块预测仅需187毫秒,相较像素空间的世界行动模型实现了高达24倍的时钟时间延迟降低。

0
下载
关闭预览

相关内容

世界动作模型: 具身AI的下一个前沿
专知会员服务
23+阅读 · 5月13日
【综述】 机器人学习中的世界模型:全面综述
专知会员服务
21+阅读 · 5月4日
《军事行动中的人机AI编队本体模型》
专知会员服务
37+阅读 · 2025年11月2日
展望:模型驱动的深度学习
人工智能学家
12+阅读 · 2018年1月23日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
51+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
VIP会员
最新内容
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 今天4:08
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
“史诗怒火”行动:现代多域作战的重要节点
专知会员服务
8+阅读 · 7月30日
《下一代无线网络中的多无人机通信资源管理》
相关基金
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
51+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
Top
微信扫码咨询专知VIP会员