Long-horizon robot mobile manipulation requires continual reasoning about localization, environment changes, and task progress, all of which are challenging to infer from image observations alone. In this paper, we show that conditioning a mobile manipulation policy on a spatiotemporal feature map improves reasoning over long horizons. The map represents the environment and the articulated robot body as neural points in a shared latent space and is updated online from egocentric observations and proprioceptive state. We update the environment neural points using object-level rigid tracking and the robot neural points using forward kinematics. We use our spatiotemporal environment and robot feature (SERF) map as a state input to a vision-language-action (VLA) model by extracting map tokens from multiple reference frames and spatial scales, providing the policy with both local and global context. We demonstrate SERF on BEHAVIOR-1K, a benchmark for long-horizon mobile manipulation in household environments. Experiments show that the SERF VLA policy outperforms image-only baselines, reaches subgoals faster by following more direct trajectories, improves robustness to scene-configuration shifts, and recovers from object-drop failures.
翻译:摘要:长时域机器人移动操作需要对定位、环境变化及任务进程进行持续推理,而这些仅凭图像观测难以准确推断。本文证明,将移动操作策略与时空特征地图相结合可提升长时域推理能力。该地图将环境与铰接式机器人本体表示为共享潜在空间中的神经点,并通过自我中心观测与本体感觉状态进行在线更新。我们利用物体级刚性追踪更新环境神经点,通过正向运动学更新机器人神经点。通过从多参考帧与空间尺度中提取地图标记,将所提出的时空环境与机器人特征(SERF)地图作为状态输入至视觉-语言-动作(VLA)模型,为策略提供局部与全局上下文。我们在家庭环境长时域移动操作基准BEHAVIOR-1K上验证SERF的性能。实验表明,SERF VLA策略优于纯图像基线方法,能通过更直接的轨迹更快达到子目标,对场景配置变化具有更强鲁棒性,并能从物体掉落故障中恢复。