Min-max problems are important in multi-agent sequential decision-making because they improve the performance of the worst-performing agent in the network. However, solving the multi-agent min-max problem is challenging. We propose a modular, distributed, online planning-based algorithm that is able to approximate the solution of the min-max objective in networked Markov games, assuming that the agents communicate within a network topology and the transition and reward functions are neighborhood-dependent. This set-up is encountered in the multi-robot setting. Our method consists of two phases at every planning step. In the first phase, each agent obtains sample returns based on its local reward function, by performing online planning. Using the samples from online planning, each agent constructs a concave approximation of its underlying local return as a function of only the action of its neighborhood at the next planning step. In the second phase, the agents deploy a distributed optimization framework that converges to the optimal immediate next action for each agent, based on the function approximations of the first phase. We demonstrate our algorithm's performance through formation control simulations.
翻译:极小极大问题在多智能体序贯决策中至关重要,因为它能提升网络中性能最差智能体的表现。然而,求解多智能体极小极大问题具有挑战性。我们提出了一种模块化、分布式、基于在线规划的算法,该算法能够逼近网络化马尔可夫博弈中极小极大目标的解,其假设智能体在网络拓扑内进行通信,且状态转移函数与奖励函数均依赖于邻域。这种设定常见于多机器人场景。我们的方法在每个规划步包含两个阶段。第一阶段,每个智能体通过执行在线规划,基于其局部奖励函数获取采样回报。利用在线规划得到的样本,每个智能体将其底层局部回报构建为一个仅关于下一规划步其邻域动作的凹函数近似。第二阶段,智能体部署一个分布式优化框架,该框架基于第一阶段的函数近似,收敛至每个智能体的最优即时下一动作。我们通过编队控制仿真验证了算法的性能。