Robots that learn over long deployments must add new skills without losing the shared policy structure that makes earlier skills reusable. We study sequential robot skill learning, where previous trajectories and task losses may be unavailable, and the deployed policy must remain a single shared controller without task-specific heads, routing, or adapters. We identify skill-coupling collapse, a failure mode in which individual skill success remains non-trivial while reliability among related skills deteriorates. We propose Sleeping Robots, a wake-sleep framework that learns each new skill during wake and consolidates the shared policy offline during sleep using compact frozen skill memories: frozen critics with unordered state buffers for reinforcement learning and frozen actor snapshots with unordered observation buffers for imitation learning. During sleep, these memories define differentiable surrogate objectives whose gradients are combined through Nash bargaining, with adaptive anchoring and local excitability for stable consolidation. On Meta-World MT5, Sleeping Robots improves average success by 64 % and pairwise reliability by x 2.0 over the strongest non-oracle baseline, and on SurgicAI it improves average success and backward transfer relative to continual imitation baselines while remaining competitive on pairwise reliability.
翻译:在长期部署中学习的机器人必须在不丢失使早期技能可复用的共享策略结构的前提下,添加新技能。我们研究序列式机器人技能学习,其中先前的轨迹和任务损失可能不可用,且部署的策略必须保持单一共享控制器,无需任务特定的头部、路由或适配器。我们识别出技能耦合崩溃这一失效模式,其中单个技能的成功率保持非平凡,但相关技能之间的可靠性却降低。我们提出"沉睡机器人"框架,这是一个"唤醒-睡眠"框架:在唤醒阶段学习每个新技能,在睡眠阶段利用紧凑的冻结技能记忆(强化学习中使用冻结的评论家与无序状态缓冲区,模仿学习中使用冻结的演员快照与无序观测缓冲区)离线巩固共享策略。在睡眠期间,这些记忆定义了可微分的代理目标,其梯度通过纳什议价进行组合,并辅以自适应锚定与局部兴奋性以实现稳定巩固。在Meta-World MT5上,"沉睡机器人"将平均成功率提升了64%,成对可靠性提升了2.0倍,优于最强非先知基线;在SurgicAI上,其相对于持续模仿学习基线提升了平均成功率和反向迁移性能,同时在成对可靠性方面保持竞争力。