Robot action planning in the real world is challenging as it requires not only understanding the current state of the environment but also predicting how it will evolve in response to actions. Vision-language-action (VLA), which repurpose large-scale vision-language models for robot action generation using action experts, have achieved notable success across a variety of robotic tasks. Nevertheless, their performance remains constrained by the scope of their training data, exhibiting limited generalization to unseen scenarios and vulnerability to diverse contextual perturbations. More recently, world models have been revisited as an alternative to VLAs. These models, referred to as world action models (WAMs), are built upon world models that are trained on large corpora of video data to predict future states. With minor adaptations, their latent representation can be decoded into robot actions. It has been suggested that their explicit dynamic prediction capacity, combined with spatiotemporal priors acquired from web-scale video pretraining, enables WAMs to generalize more effectively than VLAs. In this paper, we conduct a comparative study of prominent state-of-the-art VLA policies and recently released WAMs. We evaluate their performance on the LIBERO-Plus and RoboTwin 2.0-Plus benchmarks under various visual and language perturbations. Our results show that WAMs achieve strong robustness, with LingBot-VA reaching 74.2% success rate on RoboTwin 2.0-Plus and Cosmos-Policy achieving 82.2% on LIBERO-Plus. While VLAs such as $π_{0.5}$ can achieve comparable robustness on certain tasks, they typically require extensive training with diverse robotic datasets and varied learning objectives. Hybrid approaches that partially incorporate video-based dynamic learning exhibit intermediate robustness, highlighting the importance of how video priors are integrated.


翻译:在真实世界中进行机器人动作规划具有挑战性,这不仅需要理解环境的当前状态,还需预测其对动作的响应演化过程。视觉-语言-动作模型通过利用动作专家重新利用大规模视觉-语言模型生成机器人动作,已在多种机器人任务中取得显著成功。然而,其性能仍受限于训练数据的范围,对未见场景的泛化能力有限,且易受多种情境扰动影响。近期,世界模型作为视觉-语言-动作模型的替代方案被重新审视。这类被称为世界动作模型的系统,基于通过大规模视频数据语料库训练以预测未来状态的世界模型构建。通过微小调整,其潜在表示可解码为机器人动作。研究表明,其显式的动态预测能力结合从网络规模视频预训练中获得的时空先验,使世界动作模型比视觉-语言-动作模型更具泛化能力。本文对当前最先进的视觉-语言-动作模型策略与近期发布的世界动作模型进行了对比研究,在LIBERO-Plus和RoboTwin 2.0-Plus基准测试中评估了多种视觉与语言扰动下的性能表现。结果表明,世界动作模型展现出强鲁棒性:LingBot-VA在RoboTwin 2.0-Plus上达到74.2%的成功率,Cosmos-Policy在LIBERO-Plus上达到82.2%的成功率。尽管像$π_{0.5}$这样的视觉-语言-动作模型在某些任务中可达到相当的鲁棒性,但通常需要利用多样化机器人数据集和不同学习目标进行大量训练。部分融合视频动态学习的混合方法表现出中等鲁棒性,凸显了视频先验集成方式的重要性。

0
下载
关闭预览

相关内容

世界动作模型: 具身AI的下一个前沿
专知会员服务
24+阅读 · 5月13日
机器人领域的多任务泛化研究
专知会员服务
16+阅读 · 1月14日
视觉-语言-动作模型解析:从模块构成到里程碑与挑战
专知会员服务
17+阅读 · 2025年12月17日
面向具身操作的高效视觉–语言–动作模型:系统综述
专知会员服务
27+阅读 · 2025年10月22日
视觉-语言-动作(VLA)模型的前世今生
专知会员服务
22+阅读 · 2025年8月29日
面向具身操作的视觉-语言-动作模型综述
专知会员服务
28+阅读 · 2025年8月23日
视觉语言动作模型:概念、进展、应用与挑战
专知会员服务
19+阅读 · 2025年5月18日
深度学习模型可解释性的研究进展
专知
26+阅读 · 2020年8月1日
这可能是「多模态机器学习」最通俗易懂的介绍
计算机视觉life
113+阅读 · 2018年12月20日
展望:模型驱动的深度学习
人工智能学家
12+阅读 · 2018年1月23日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
10+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
51+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
VIP会员
最新内容
机器的崛起:美海军陆战队组建机器人营思考
专知会员服务
6+阅读 · 9月8日
无面之战:人工智能如何重绘权力版图
专知会员服务
4+阅读 · 9月8日
《最强大的军事网状网络》
专知会员服务
9+阅读 · 9月7日
《预测陆军征兵任务分配》110页
专知会员服务
7+阅读 · 9月7日
相关VIP内容
世界动作模型: 具身AI的下一个前沿
专知会员服务
24+阅读 · 5月13日
机器人领域的多任务泛化研究
专知会员服务
16+阅读 · 1月14日
视觉-语言-动作模型解析:从模块构成到里程碑与挑战
专知会员服务
17+阅读 · 2025年12月17日
面向具身操作的高效视觉–语言–动作模型:系统综述
专知会员服务
27+阅读 · 2025年10月22日
视觉-语言-动作(VLA)模型的前世今生
专知会员服务
22+阅读 · 2025年8月29日
面向具身操作的视觉-语言-动作模型综述
专知会员服务
28+阅读 · 2025年8月23日
视觉语言动作模型:概念、进展、应用与挑战
专知会员服务
19+阅读 · 2025年5月18日
相关基金
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
10+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
51+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
Top
微信扫码咨询专知VIP会员