We present AEM, an Action-Effect Memory pretraining framework for robot manipulation that learns compact temporal representations from vision-action history. Unlike prior robot representation pretraining methods that mainly focus on single-frame visual encoding, AEM targets the temporal nature of manipulation, where the current observation alone is often insufficient under partial observability. AEM models manipulation as an action-driven interaction process by interleaving visual and action features and applying masked modeling to recover missing content from incomplete histories, thereby learning action-conditioned state evolution. The Mamba-encoded output of the final vision token is used as a compact history representation, serving as the global context for decoding and downstream control. This design preserves a single-vector temporal bottleneck while keeping inference efficient. We evaluate AEM with Diffusion Policy and Flow Policy. AEM consistently improves manipulation performance in both simulation and real-world settings, outperforming baselines across clean scenes, cluttered and random scenes, and non-Markovian tasks. Ablation studies further show that history-aware pretraining surpasses single-frame pretraining and direct frame stacking, while reducing inference latency and computational cost.
翻译:我们提出了AEM——一种面向机器人操作的效应记忆预训练框架,该框架能从视觉-动作历史中学习紧凑的时序表示。与以往主要关注单帧视觉编码的机器人表示预训练方法不同,AEM聚焦于操作任务的时序特性——在部分可观测条件下,仅凭当前观测往往不足以完成任务。AEM通过交替排列视觉与动作特征将操作建模为动作驱动的交互过程,并应用掩码建模从残缺历史中恢复缺失内容,从而学习动作条件化的状态演化。最终视觉token经Mamba编码的输出被用作紧凑的历史表示,为解码与下游控制提供全局上下文。该设计在保持高效推理的同时,保留了单向量时序瓶颈机制。我们在扩散策略与流程策略上评估了AEM。在仿真与真实世界场景中,AEM持续提升操作性能,在整洁场景、杂乱随机场景及非马尔可夫任务中均优于基线方法。消融实验进一步表明,历史感知预训练优于单帧预训练与直接帧拼接,同时降低了推理延迟与计算开销。