Decentralized partially observable Markov decision processes (Dec-POMDPs) formalize the problem of designing individual controllers for a group of collaborative agents under stochastic dynamics and partial observability. Seeking a global optimum is difficult (NEXP complete), but seeking a Nash equilibrium -- each agent policy being a best response to the other agents -- is more accessible, and allowed addressing infinite-horizon problems with solutions in the form of finite state controllers. In this paper, we show that this approach can be adapted to cases where only a generative model (a simulator) of the Dec-POMDP is available. This requires relying on a simulation-based POMDP solver to construct an agent's FSC node by node. A related process is used to heuristically derive initial FSCs. Experiment with benchmarks shows that MC-JESP is competitive with exisiting Dec-POMDP solvers, even better than many offline methods using explicit models.
翻译:去中心化部分可观测马尔可夫决策过程(Dec-POMDPs)形式化了在随机动力学与部分可观测条件下为一组协作智能体设计独立控制器的问题。寻找全局最优解是困难的(NEXP完全问题),但寻找纳什均衡——每个智能体策略均为对其他智能体策略的最优响应——则更为可行,并允许以有限状态控制器形式解决无限时域问题。本文表明,该方法可适用于仅拥有Dec-POMDP生成模型(模拟器)的情形。这需要借助基于模拟的POMDP求解器逐节点构建智能体的FSC。相关过程用于启发式地推导初始FSC。基准实验表明,MC-JESP与现有Dec-POMDP求解器竞争力相当,甚至优于许多使用显式模型的离线方法。