Multi-Agent Policy Gradient (MAPG) has made significant progress in recent years. However, centralized critics in state-of-the-art MAPG methods still face the centralized-decentralized mismatch (CDM) issue, which means sub-optimal actions by some agents will affect other agent's policy learning. While using individual critics for policy updates can avoid this issue, they severely limit cooperation among agents. To address this issue, we propose an agent topology framework, which decides whether other agents should be considered in policy gradient and achieves compromise between facilitating cooperation and alleviating the CDM issue. The agent topology allows agents to use coalition utility as learning objective instead of global utility by centralized critics or local utility by individual critics. To constitute the agent topology, various models are studied. We propose Topology-based multi-Agent Policy gradiEnt (TAPE) for both stochastic and deterministic MAPG methods. We prove the policy improvement theorem for stochastic TAPE and give a theoretical explanation for the improved cooperation among agents. Experiment results on several benchmarks show the agent topology is able to facilitate agent cooperation and alleviate CDM issue respectively to improve performance of TAPE. Finally, multiple ablation studies and a heuristic graph search algorithm are devised to show the efficacy of the agent topology.
翻译:多智能体策略梯度(MAPG)近年来取得了显著进展。然而,当前最先进的MAPG方法中的集中式评论家仍面临集中-分散不匹配(CDM)问题,即部分智能体的次优行为会影响其他智能体的策略学习。虽然采用个体评论家进行策略更新可避免该问题,但会严重限制智能体间的协作。为应对这一挑战,我们提出了一种智能体拓扑框架,该框架可决定策略梯度中是否考虑其他智能体,并在促进协作与缓解CDM问题之间取得折中。智能体拓扑允许智能体以联盟效用而非集中式评论家的全局效用或个体评论家的局部效用作为学习目标。为构建智能体拓扑,我们研究了多种模型,并提出了适用于随机和确定性MAPG方法的拓扑多智能体策略梯度(TAPE)。我们证明了随机TAPE的策略提升定理,并从理论上解释了智能体间协作增强的原因。在多个基准测试上的实验结果表明,智能体拓扑能分别促进协作与缓解CDM问题,从而提升TAPE性能。最后,通过多项消融实验与启发式图搜索算法验证了智能体拓扑的有效性。