The multiple-target self-organizing pursuit (SOP) problem has wide applications and has been considered a challenging self-organization game for distributed systems, in which intelligent agents cooperatively pursue multiple dynamic targets with partial observations. This work proposes a framework for decentralized multi-agent systems to improve the implicit coordination capabilities in search and pursuit. We model a self-organizing system as a partially observable Markov game (POMG) featured by large-scale, decentralization, partial observation, and noncommunication. The proposed distributed algorithm: fuzzy self-organizing cooperative coevolution (FSC2) is then leveraged to resolve the three challenges in multi-target SOP: distributed self-organizing search (SOS), distributed task allocation, and distributed single-target pursuit. FSC2 includes a coordinated multi-agent deep reinforcement learning (MARL) method that enables homogeneous agents to learn natural SOS patterns. Additionally, we propose a fuzzy-based distributed task allocation method, which locally decomposes multi-target SOP into several single-target pursuit problems. The cooperative coevolution principle is employed to coordinate distributed pursuers for each single-target pursuit problem. Therefore, the uncertainties of inherent partial observation and distributed decision-making in the POMG can be alleviated. The experimental results demonstrate that by decomposing the SOP task, FSC2 achieves superior performance compared with other implicit coordination policies fully trained by general MARL algorithms. The scalability of FSC2 is proved that up to 2048 FSC2 agents perform efficient multi-target SOP with almost 100 percent capture rates. Empirical analyses and ablation studies verify the interpretability, rationality, and effectiveness of component algorithms in FSC2.
翻译:多目标自组织追捕(SOP)问题具有广泛应用,被视为分布式系统中具有挑战性的自组织博弈问题,其中智能体需在部分观测条件下协作追捕多个动态目标。本文提出一种去中心化多智能体系统框架,以提升搜索与追捕过程中的隐式协调能力。我们将自组织系统建模为具有大规模、去中心化、部分观测和无通信特征的部分可观测马尔可夫博弈(POMG),进而采用所提出的分布式算法——模糊自组织协作协同进化(FSC2)解决多目标SOP中的三大挑战:分布式自组织搜索(SOS)、分布式任务分配和分布式单目标追捕。FSC2包含一种协调式多智能体深度强化学习(MARL)方法,使同构智能体能够学习自然的SOS模式。此外,我们提出一种基于模糊的分布式任务分配方法,将多目标SOP局部分解为多个单目标追捕问题。通过引入协作协同进化原理,协调各单目标追捕问题的分布式追捕者,从而缓解POMG中固有部分观测与分布式决策带来的不确定性。实验结果表明,通过分解SOP任务,与通用MARL算法完全训练的其他隐式协调策略相比,FSC2展现出更优越的性能。FSC2的可扩展性得以验证:多达2048个FSC2智能体可实现高效的多目标SOP,捕获率接近100%。实证分析与消融实验验证了FSC2各组成算法的可解释性、合理性与有效性。