Fleets of robots ingest massive amounts of streaming data generated by interacting with their environments, far more than those that can be stored or transmitted with ease. At the same time, we hope that teams of robots can co-acquire diverse skills through their experiences in varied settings. How can we enable such fleet-level learning without having to transmit or centralize fleet-scale data? In this paper, we investigate distributed learning of policies as a potential solution. To efficiently merge policies in the distributed setting, we propose fleet-merge, an instantiation of distributed learning that accounts for the symmetries that can arise in learning policies that are parameterized by recurrent neural networks. We show that fleet-merge consolidates the behavior of policies trained on 50 tasks in the Meta-World environment, with the merged policy achieving good performance on nearly all training tasks at test time. Moreover, we introduce a novel robotic tool-use benchmark, fleet-tools, for fleet policy learning in compositional and contact-rich robot manipulation tasks, which might be of broader interest, and validate the efficacy of fleet-merge on the benchmark.
翻译:机器人群在与其环境交互过程中产生海量流式数据,其规模远超可便捷存储或传输的限度。与此同时,我们期望机器人团队能通过在不同环境中的经历共同习得多样化技能。如何在无需传输或集中化集群规模数据的前提下实现这种集群级学习?本文研究将分布式策略学习作为潜在解决方案。为在分布式场景下高效合并策略,我们提出fleet-merge方法——一种考虑循环神经网络参数化策略学习过程中可能出现的对称性的分布式学习实例。实验表明,fleet-merge能够整合在Meta-World环境中50个任务上训练的机器人策略行为,合并后的策略在测试时几乎对所有训练任务均表现优异。此外,我们引入了一个新颖的机器人工具使用基准fleet-tools(可能具有更广泛的研究价值),用于组合式与高接触性机器人操控任务中的集群策略学习,并在该基准上验证了fleet-merge的有效性。