World Action Models (WAMs) offer a promising route for robot manipulation by using video generation models to model future scene evolution before producing control actions. However, our empirical observations reveal a phenomenon: generating plausible visual futures does not always guarantee the extraction of accurate actions. To diagnose this failure, we conduct action-head attention analysis and causal interventions. We find that the action decoder fails to focus on task-relevant interaction regions and remains sensitive to perturbations in task-irrelevant areas. This reveals a representation mismatch: hidden states optimized for visual reconstruction are not inherently organized in a form useful for low-level action control. In this paper, we propose AGRA, an Action-Grounded Representation Alignment objective that regularizes the world-action interface by aligning intermediate video diffusion features with spatially coherent semantic representations from a foundation visual encoder. We evaluate AGRA on real-world manipulation tasks. Experiments show that AGRA makes world model representations more action-grounded: by focusing the action decoder on the correct interaction regions, it improves object localization accuracy and affordance understanding, and makes the policy more robust to perturbations in task-irrelevant regions. As a result, AGRA consistently improves both in-distribution performance and out-of-distribution generalization over the baseline world action model.
翻译:世界动作模型通过使用视频生成模型在生成控制动作之前建模未来场景演变,为机器人操作提供了一条有前景的路径。然而,我们的实证观察揭示了一个现象:生成合理的视觉未来并不总能保证提取准确的动作。为了诊断这一失败,我们进行了动作头部注意力分析和因果干预。我们发现,动作解码器未能聚焦于任务相关的交互区域,并且对任务无关区域的扰动敏感。这揭示了一种表示不匹配:为视觉重建优化的隐藏状态并未固有地以适用于低级动作控制的形式组织。在本文中,我们提出了AGRA,一种动作接地表示对齐目标,通过将中间视频扩散特征与基础视觉编码器的空间一致语义表示对齐,来正则化世界动作接口。我们在真实世界的操作任务上评估了AGRA。实验表明,AGRA使世界模型表示更具动作接地性:通过将动作解码器聚焦于正确的交互区域,它提高了对象定位准确性和功能理解,并使策略对任务无关区域的扰动更加鲁棒。因此,AGRA在分布内性能和分布外泛化能力上均始终优于基线世界动作模型。