Hand-eye calibration through visual localization is a critical capability for robotic manipulation in open-world environments. However, most deep learning-based calibration models suffer from catastrophic forgetting when adapting into unseen data amongst open-world scene changes, while simple rehearsal-based continual learning strategy cannot well mitigate this issue. To overcome this challenge, we propose a continual hand-eye calibration framework, enabling robots to adapt to sequentially encountered open-world manipulation scenes through spatially replay strategy and structure-preserving distillation. Specifically, a Spatial-Aware Replay Strategy (SARS) constructs a geometrically uniform replay buffer that ensures comprehensive coverage of each scene pose space, replacing redundant adjacent frames with maximally informative viewpoints. Meanwhile, a Structure-Preserving Dual Distillation (SPDD) is proposed to decompose localization knowledge into coarse scene layout and fine pose precision, and distills them separately to alleviate both types of forgetting during continual adaptation. As a new manipulation scene arrives, SARS provides geometrically representative replay samples from all prior scenes, and SPDD applies structured distillation on these samples to retain previously learned knowledge. After training on the new scene, SARS incorporates selected samples from the new scene into the replay buffer for future rehearsal, allowing the model to continuously accumulate multi-scene calibration capability. Experiments on multiple public datasets show significant anti scene forgetting performance, maintaining accuracy on past scenes while preserving adaptation to new scenes, confirming the effectiveness of the framework.
翻译:通过视觉定位进行手眼标定是开放世界环境中机器人操作的一项关键能力。然而,当面对开放世界场景变化中的未见数据时,多数基于深度学习的标定模型会出现灾难性遗忘,而简单的基于回放的持续学习策略难以有效缓解该问题。为克服这一挑战,我们提出了一种持续手眼标定框架,通过空间回放策略与保结构蒸馏,使机器人能适应序列出现的开放世界操作场景。具体而言,空间感知回放策略(SARS)构建了一个几何均匀的回放缓存,确保覆盖每个场景位姿空间,用信息量最大的视角替代冗余的相邻帧。同时,我们提出了保结构双重蒸馏(SPDD)方法,将定位知识分解为粗略场景布局与精细位姿精度,并分别蒸馏以缓解持续适应过程中两种类型的遗忘。当新操作场景出现时,SARS从所有历史场景中提取具几何代表性的回放样本,SPDD则对这些样本应用结构化蒸馏以保留先前学到的知识。在新场景训练完成后,SARS将新场景中选取的样本纳入回放缓存供未来回放,使模型能持续累积多场景标定能力。在多个公开数据集上的实验显示出显著的抗场景遗忘性能,在保持对历史场景精度的同时维持对新场景的适应能力,验证了该框架的有效性。