Role-based learning is a promising approach to improving the performance of Multi-Agent Reinforcement Learning (MARL). Nevertheless, without manual assistance, current role-based methods cannot guarantee stably discovering a set of roles to effectively decompose a complex task, as they assume either a predefined role structure or practical experience for selecting hyperparameters. In this article, we propose a mathematical Structural Information principles-based Role Discovery method, namely SIRD, and then present a SIRD optimizing MARL framework, namely SR-MARL, for multi-agent collaboration. The SIRD transforms role discovery into a hierarchical action space clustering. Specifically, the SIRD consists of structuralization, sparsification, and optimization modules, where an optimal encoding tree is generated to perform abstracting to discover roles. The SIRD is agnostic to specific MARL algorithms and flexibly integrated with various value function factorization approaches. Empirical evaluations on the StarCraft II micromanagement benchmark demonstrate that, compared with state-of-the-art MARL algorithms, the SR-MARL framework improves the average test win rate by 0.17%, 6.08%, and 3.24%, and reduces the deviation by 16.67%, 30.80%, and 66.30%, under easy, hard, and super hard scenarios.
翻译:基于角色的学习是提升多智能体强化学习性能的一种有前景的方法。然而,在没有人工辅助的情况下,当前的基于角色的方法无法保证稳定地发现一组角色来有效分解复杂任务,因为它们要么预设了角色结构,要么依赖实践经验选择超参数。本文提出了一种基于数学结构信息原理的角色发现方法SIRD,并在此基础上提出了一个优化多智能体强化学习框架SR-MARL,用于多智能体协作。SIRD将角色发现转化为层次化动作空间聚类问题。具体来说,SIRD包含结构化、稀疏化和优化三个模块,通过生成最优编码树进行抽象以发现角色。SIRD与特定多智能体强化学习算法无关,可灵活集成各种值函数分解方法。在星际争霸II微观管理基准测试上的实证评估表明,与最先进的多智能体强化学习算法相比,SR-MARL框架在简单、困难与极困难场景下的平均测试胜率分别提升了0.17%、6.08%与3.24%,偏差则分别降低了16.67%、30.80%与66.30%。