We describe a system for deep reinforcement learning of robotic manipulation skills applied to a large-scale real-world task: sorting recyclables and trash in office buildings. Real-world deployment of deep RL policies requires not only effective training algorithms, but the ability to bootstrap real-world training and enable broad generalization. To this end, our system combines scalable deep RL from real-world data with bootstrapping from training in simulation, and incorporates auxiliary inputs from existing computer vision systems as a way to boost generalization to novel objects, while retaining the benefits of end-to-end training. We analyze the tradeoffs of different design decisions in our system, and present a large-scale empirical validation that includes training on real-world data gathered over the course of 24 months of experimentation, across a fleet of 23 robots in three office buildings, with a total training set of 9527 hours of robotic experience. Our final validation also consists of 4800 evaluation trials across 240 waste station configurations, in order to evaluate in detail the impact of the design decisions in our system, the scaling effects of including more real-world data, and the performance of the method on novel objects. The projects website and videos can be found at \href{http://rl-at-scale.github.io}{rl-at-scale.github.io}.
翻译:我们描述了一套用于机器人操作技能的深度强化学习系统,该系统应用于大规模真实世界任务:在办公建筑中分拣可回收物和垃圾。深度强化学习策略在真实世界的部署不仅需要有效的训练算法,还需具备启动真实世界训练并实现广泛泛化的能力。为此,我们的系统结合了基于真实世界数据的可扩展深度强化学习与仿真训练的启动能力,并整合来自现有计算机视觉系统的辅助输入,以提升对新颖物体的泛化能力,同时保留端到端训练的优势。我们分析了系统中不同设计决策的权衡,并进行了大规模实证验证,包括在24个月的实验期间收集的真实世界数据上进行训练,涉及三栋办公建筑中的23台机器人组成的车队,总计9527小时的机器人经验数据。最终验证还包括在240个垃圾站配置上的4800次评估试验,以详细评估系统设计决策的影响、包含更多真实世界数据的扩展效应,以及该方法在新颖物体上的性能。项目网站和视频可在 \href{http://rl-at-scale.github.io}{rl-at-scale.github.io} 查看。