World Action Models (WAMs) have emerged as a new powerful paradigm for embodied intelligence, learning action-relevant visual dynamics that significantly enhance generalization and robustness. However, existing WAMs still struggle with task-relevant memory in long-horizon robotic manipulation. To address this, we present HiMem-WAM, a Hierarchical Memory-Gated WAM that integrates motion-centric latent actions, high-level skill latents, and boundary-triggered memory updates. Specifically, we develop a hierarchical latent action framework that jointly learns low-level motion and high-level skill latents, providing structured temporal abstraction. Meanwhile, a boundary-aware memory gate writes compact task states at predicted skill transitions, enabling causal inference without test-time generation of future video or optical flow estimation. Evaluated on LIBERO, LIBERO-PLUS, RMBench and real-world tasks, HiMem-WAM shows that hierarchical latents improve robustness under deployment perturbations, and the memory module substantially benefits memory-dependent long-horizon manipulation.
翻译:世界动作模型(WAMs)已成为具身智能领域一种强大的新范式,通过学习与动作相关的视觉动态,显著提升了泛化能力与鲁棒性。然而,现有WAMs在长时域机器人操作中仍面临与任务相关记忆的挑战。为此,我们提出HiMem-WAM,一种集成动作中心隐变量、高层技能隐变量以及边界触发记忆更新的分层记忆门控WAM。具体而言,我们构建了一个分层隐变量动作框架,联合学习低层运动隐变量与高层技能隐变量,实现结构化的时间抽象。同时,边界感知记忆门在预测的技能转换点写入紧凑的任务状态,从而实现无需测试时生成未来视频或光流估计的因果推理。在LIBERO、LIBERO-PLUS、RMBench及真实世界任务上的评估表明,HiMem-WAM的分层隐变量提升了部署扰动下的鲁棒性,而记忆模块则显著有利于依赖记忆的长时域操作。