In centralized multi-agent systems, often modeled as multi-agent partially observable Markov decision processes (MPOMDPs), the action and observation spaces grow exponentially with the number of agents, making the value and belief state estimation of single-agent online planning ineffective. Prior work partially tackles value estimation by exploiting the inherent structure of multi-agent settings via so-called coordination graphs. Additionally, belief state estimation has been improved by incorporating the likelihood of observations into the approximation. However, the challenges of value estimation and state estimation have only been tackled individually, which prevents these methods from scaling to many agents. Therefore, we address these challenges simultaneously. First, we introduce weighted particle filtering to sample-based online planners in MPOMDPs. Second, we present a scalable approximation of the belief state. Third, we bring an approach that exploits the typical locality of agent interactions to novel online planning algorithms for MPOMDPs operating on a so-called sparse particle filter belief tree. Our algorithms show competitive performance for settings with only a few agents and outperform state-of-the-art algorithms on benchmarks with many agents.
翻译:在集中式多智能体系统中(通常建模为多智能体部分可观测马尔可夫决策过程,即MPOMDPs),动作空间和观测空间随智能体数量呈指数级增长,这使得单智能体在线规划中的价值和信念状态估计方法失效。现有工作通过利用所谓协调图挖掘多智能体场景的固有结构,部分解决了价值估计问题。同时,通过将观测似然性纳入近似框架,信念状态估计也得到了改进。然而,价值估计与状态估计的挑战始终被单独处理,这阻碍了这些方法扩展到大规模智能体场景。为此,我们同步应对这些挑战。首先,我们为基于采样的MPOMDP在线规划器引入加权粒子滤波;其次,提出信念状态的可扩展近似方法;第三,提出一种利用智能体交互典型局部性的新方法,并基于所谓的稀疏粒子滤波信念树构建新型MPOMDP在线规划算法。我们的算法在少量智能体场景中展现出竞争性能,并在大规模智能体基准测试中显著超越现有最优算法。