We consider the multi-agent spatial navigation problem of computing the socially optimal order of play, i.e., the sequence in which the agents commit to their decisions, and its associated equilibrium in an N-player Stackelberg trajectory game. We model this problem as a mixed-integer optimization problem over the space of all possible Stackelberg games associated with the order of play's permutations. To solve the problem, we introduce Branch and Play (B&P), an efficient and exact algorithm that provably converges to a socially optimal order of play and its Stackelberg equilibrium. As a subroutine for B&P, we employ and extend sequential trajectory planning, i.e., a popular multi-agent control approach, to scalably compute valid local Stackelberg equilibria for any given order of play. We demonstrate the practical utility of B&P to coordinate air traffic control, swarm formation, and delivery vehicle fleets. We find that B&P consistently outperforms various baselines, and computes the socially optimal equilibrium.
翻译:我们考虑多智能体空间导航问题中社会最优博弈顺序的计算,即智能体承诺决策的先后顺序及其在N人Stackelberg轨迹博弈中的相关均衡。我们将该问题建模为混合整数优化问题,其解空间涵盖所有与博弈顺序排列相关的可能Stackelberg博弈。为解决该问题,我们提出分支博弈算法(Branch and Play, B&P),这是一种能够严格收敛至社会最优博弈顺序及其Stackelberg均衡的高效精确算法。作为B&P的子程序,我们采用并扩展了序列轨迹规划(一种流行的多智能体控制方法),可对任意给定博弈顺序可扩展地计算有效局部Stackelberg均衡。我们展示了B&P在协调空中交通管制、集群编队及物流车队等方面的实际效用。实验结果表明,B&P始终优于多种基线方法,并能计算社会最优均衡。