Robot action planning in the real world is challenging as it requires not only understanding the current state of the environment but also predicting how it will evolve in response to actions. Vision-language-action (VLA), which repurpose large-scale vision-language models for robot action generation using action experts, have achieved notable success across a variety of robotic tasks. Nevertheless, their performance remains constrained by the scope of their training data, exhibiting limited generalization to unseen scenarios and vulnerability to diverse contextual perturbations. More recently, world models have been revisited as an alternative to VLAs. These models, referred to as world action models (WAMs), are built upon world models that are trained on large corpora of video data to predict future states. With minor adaptations, their latent representation can be decoded into robot actions. It has been suggested that their explicit dynamic prediction capacity, combined with spatiotemporal priors acquired from web-scale video pretraining, enables WAMs to generalize more effectively than VLAs. In this paper, we conduct a comparative study of prominent state-of-the-art VLA policies and recently released WAMs. We evaluate their performance on the LIBERO-Plus and RoboTwin 2.0-Plus benchmarks under various visual and language perturbations. Our results show that WAMs achieve strong robustness, with LingBot-VA reaching 74.2% success rate on RoboTwin 2.0-Plus and Cosmos-Policy achieving 82.2% on LIBERO-Plus. While VLAs such as $π_{0.5}$ can achieve comparable robustness on certain tasks, they typically require extensive training with diverse robotic datasets and varied learning objectives. Hybrid approaches that partially incorporate video-based dynamic learning exhibit intermediate robustness, highlighting the importance of how video priors are integrated.
翻译:在真实世界中进行机器人动作规划具有挑战性,这不仅需要理解环境的当前状态,还需预测其对动作的响应演化过程。视觉-语言-动作模型通过利用动作专家重新利用大规模视觉-语言模型生成机器人动作,已在多种机器人任务中取得显著成功。然而,其性能仍受限于训练数据的范围,对未见场景的泛化能力有限,且易受多种情境扰动影响。近期,世界模型作为视觉-语言-动作模型的替代方案被重新审视。这类被称为世界动作模型的系统,基于通过大规模视频数据语料库训练以预测未来状态的世界模型构建。通过微小调整,其潜在表示可解码为机器人动作。研究表明,其显式的动态预测能力结合从网络规模视频预训练中获得的时空先验,使世界动作模型比视觉-语言-动作模型更具泛化能力。本文对当前最先进的视觉-语言-动作模型策略与近期发布的世界动作模型进行了对比研究,在LIBERO-Plus和RoboTwin 2.0-Plus基准测试中评估了多种视觉与语言扰动下的性能表现。结果表明,世界动作模型展现出强鲁棒性:LingBot-VA在RoboTwin 2.0-Plus上达到74.2%的成功率,Cosmos-Policy在LIBERO-Plus上达到82.2%的成功率。尽管像$π_{0.5}$这样的视觉-语言-动作模型在某些任务中可达到相当的鲁棒性,但通常需要利用多样化机器人数据集和不同学习目标进行大量训练。部分融合视频动态学习的混合方法表现出中等鲁棒性,凸显了视频先验集成方式的重要性。