Robots that learn over long deployments must add new skills without losing the shared policy structure that makes earlier skills reusable. We study sequential robot skill learning, where previous trajectories and task losses may be unavailable, and the deployed policy must remain a single shared controller without task-specific heads, routing, or adapters. We identify skill-coupling collapse, a failure mode in which individual skill success remains non-trivial while reliability among related skills deteriorates. We propose Sleeping Robots, a wake-sleep framework that learns each new skill during wake and consolidates the shared policy offline during sleep using compact frozen skill memories: frozen critics with unordered state buffers for reinforcement learning and frozen actor snapshots with unordered observation buffers for imitation learning. During sleep, these memories define differentiable surrogate objectives whose gradients are combined through Nash bargaining, with adaptive anchoring and local excitability for stable consolidation. On Meta-World MT5, Sleeping Robots improves average success by 64 % and pairwise reliability by x 2.0 over the strongest non-oracle baseline, and on SurgicAI it improves average success and backward transfer relative to continual imitation baselines while remaining competitive on pairwise reliability.


翻译:在长期部署中学习的机器人必须在不丢失使早期技能可复用的共享策略结构的前提下,添加新技能。我们研究序列式机器人技能学习,其中先前的轨迹和任务损失可能不可用,且部署的策略必须保持单一共享控制器,无需任务特定的头部、路由或适配器。我们识别出技能耦合崩溃这一失效模式,其中单个技能的成功率保持非平凡,但相关技能之间的可靠性却降低。我们提出"沉睡机器人"框架,这是一个"唤醒-睡眠"框架:在唤醒阶段学习每个新技能,在睡眠阶段利用紧凑的冻结技能记忆(强化学习中使用冻结的评论家与无序状态缓冲区,模仿学习中使用冻结的演员快照与无序观测缓冲区)离线巩固共享策略。在睡眠期间,这些记忆定义了可微分的代理目标,其梯度通过纳什议价进行组合,并辅以自适应锚定与局部兴奋性以实现稳定巩固。在Meta-World MT5上,"沉睡机器人"将平均成功率提升了64%,成对可靠性提升了2.0倍,优于最强非先知基线;在SurgicAI上,其相对于持续模仿学习基线提升了平均成功率和反向迁移性能,同时在成对可靠性方面保持竞争力。

0
下载
关闭预览

相关内容

机器人(英语:Robot)包括一切模拟人类行为或思想与模拟其他生物的机械(如机器狗,机器猫等)。狭义上对机器人的定义还有很多分类法及争议,有些电脑程序甚至也被称为机器人。在当代工业中,机器人指能自动运行任务的人造机器设备,用以取代或协助人类工作,一般会是机电设备,由计算机程序或是电子电路控制。

知识荟萃

精品入门和进阶教程、论文和代码整理等

更多

查看相关VIP内容、论文、资讯等
【斯坦福博士论文】协作多机器人学习算法
专知会员服务
17+阅读 · 2025年1月6日
干货 | 可解释的机器学习
AI科技评论
20+阅读 · 2019年7月3日
机器学习笔试题精选
人工智能头条
13+阅读 · 2018年7月22日
国家自然科学基金
15+阅读 · 2016年12月31日
国家自然科学基金
11+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
11+阅读 · 2013年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
VIP会员
最新内容
失去控制的指挥:人工智能时代的任务式指挥
专知会员服务
8+阅读 · 9月11日
美国的新国家安全科技战略思考
专知会员服务
4+阅读 · 9月11日
综述 | 面向大模型智能体的图结构个性化记忆
专知会员服务
7+阅读 · 9月10日
相关基金
国家自然科学基金
15+阅读 · 2016年12月31日
国家自然科学基金
11+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
11+阅读 · 2013年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
Top
微信扫码咨询专知VIP会员