Due to the complex interactions between agents, learning multi-agent control policy often requires a prohibited amount of data. This paper aims to enable multi-agent systems to effectively utilize past memories to adapt to novel collaborative tasks in a data-efficient fashion. We propose the Multi-Agent Coordination Skill Database, a repository for storing a collection of coordinated behaviors associated with key vectors distinctive to them. Our Transformer-based skill encoder effectively captures spatio-temporal interactions that contribute to coordination and provides a unique skill representation for each coordinated behavior. By leveraging only a small number of demonstrations of the target task, the database enables us to train the policy using a dataset augmented with the retrieved demonstrations. Experimental evaluations demonstrate that our method achieves a significantly higher success rate in push manipulation tasks compared with baseline methods like few-shot imitation learning. Furthermore, we validate the effectiveness of our retrieve-and-learn framework in a real environment using a team of wheeled robots.
翻译:由于智能体间的复杂交互,学习多智能体控制策略通常需要海量数据。本文旨在使多智能体系统能够高效利用过往经验,以数据高效的方式适应新型协作任务。我们提出多智能体协作技能数据库,该存储库用于保存一系列与独特特征向量相关联的协调行为。基于Transformer的技能编码器能有效捕捉促成协调行为的时空交互,并为每个协调行为提供独特的技能表征。通过仅利用目标任务的少量演示样本,该数据库可支持使用检索增强的数据集训练策略。实验评估表明,与少量样本模仿学习等基线方法相比,本方法在推操作任务中实现了显著更高的成功率。此外,我们使用轮式机器人团队在真实环境中验证了该"检索-学习"框架的有效性。