Standard cooperative multi-agent reinforcement learning (MARL) methods aim to find the optimal team cooperative policy to complete a task. However there may exist multiple different ways of cooperating, which usually are very needed by domain experts. Therefore, identifying a set of significantly different policies can alleviate the task complexity for them. Unfortunately, there is a general lack of effective policy diversity approaches specifically designed for the multi-agent domain. In this work, we propose a method called Moment-Matching Policy Diversity to alleviate this problem. This method can generate different team policies to varying degrees by formalizing the difference between team policies as the difference in actions of selected agents in different policies. Theoretically, we show that our method is a simple way to implement a constrained optimization problem that regularizes the difference between two trajectory distributions by using the maximum mean discrepancy. The effectiveness of our approach is demonstrated on a challenging team-based shooter.
翻译:标准合作式多智能体强化学习方法旨在寻找最优的团队合作策略以完成任务。然而,通常存在多种不同的合作方式,这恰恰是领域专家所迫切需要的。因此,识别一组显著不同的策略可以降低他们的任务复杂度。遗憾的是,目前专门针对多智能体领域设计的有效策略多样性方法普遍缺乏。在本工作中,我们提出了一种名为"矩匹配策略多样性"的方法来缓解这一问题。该方法通过将团队策略差异形式化为不同策略中选定智能体动作的差异,能够生成不同程度的多样化团队策略。理论上,我们证明了该方法是通过最大均值差异正则化两个轨迹分布差异的约束优化问题的一种简洁实现方式。我们的方法在一个具有挑战性的团队射击游戏中得到了有效性验证。