Robots solving generalist tasks need to be able to ground instructions in their past experience, since humans may refer to notable past events when giving a task (e.g., ``Take me to where the chemical spill happened yesterday''). Since memory limits make storing all past events infeasible, long-term robot memory must be selective, ideally retaining only those episodes with high utility for future tasks. However, future tasks are not typically given a priori for generalist robots. To select generically useful memories, we propose Bayesian surprise as a gating mechanism for memory formation. We present an approach to compute surprise in a semantically rich deployment-agnostic latent space provided by V-JEPA-2. Using our gated episodic memory to augment 4D scene graph-based spatial memory, we show a consistent improvement over state-of-the-art benchmarks in robot question answering, outperforming prior robot memory methods by $\geq12\%$ for temporal, spatial, and binary questions, and surpassing the performance of supervised and non-causal methods with an unsupervised causal method in event segmentation tasks.
翻译:解决通用任务的机器人需要能够将指令锚定在其过往经验中,因为人类在布置任务时可能会提及显著的历史事件(例如:"带我去昨天发生化学品泄漏的地方")。由于记忆容量限制使得存储所有历史事件不可行,长期机器人记忆必须具备选择性,理想情况下仅保留那些对未来任务具有高实用价值的情景。然而,对于通用机器人而言,未来任务通常无法预先设定。为了选择具有通用价值的记忆,我们提出将贝叶斯惊奇作为记忆形成的门控机制。我们提出了一种方法,在V-JEPA-2提供的语义丰富且部署环境无关的潜在空间中计算惊奇值。通过将所提出的门控情景记忆与基于4D场景图的增强空间记忆相结合,我们在机器人问答任务的基准测试中始终优于现有最优方法:在时间、空间和二元三类问题上,性能较现有机器人记忆方法提升≥12%;在事件分割任务中,采用无监督因果方法的表现甚至超越了有监督方法和非因果方法。