In this paper, we extend mean-field Langevin dynamics to minimax optimization over probability distributions for the first time with symmetric and provably convergent updates. We propose mean-field Langevin averaged gradient (MFL-AG), a single-loop algorithm that implements gradient descent ascent in the distribution spaces with a novel weighted averaging, and establish average-iterate convergence to the mixed Nash equilibrium. We also study both time and particle discretization regimes and prove a new uniform-in-time propagation of chaos result which accounts for the dependency of the particle interactions on all previous distributions. Furthermore, we propose mean-field Langevin anchored best response (MFL-ABR), a symmetric double-loop algorithm based on best response dynamics with linear last-iterate convergence. Finally, we study applications to zero-sum Markov games and conduct simulations demonstrating long-term optimality.
翻译:本文首次将平均场朗之万动力学扩展到概率分布上的极小极大优化,并提出了具有对称性和可证明收敛性的更新方法。我们提出平均场朗之万平均梯度(MFL-AG),这是一种在分布空间中实现梯度下降上升的单循环算法,采用新颖的加权平均方法,并建立了平均迭代收敛到混合纳什均衡的理论。同时,我们研究了时间和粒子离散化机制,并证明了一个新的关于传播混沌的一致时间性结果,该结果考虑了粒子相互作用对所有先前分布的依赖性。此外,我们提出平均场朗之万锚定最优响应(MFL-ABR),这是一种基于最优响应动力学且具有线性最后迭代收敛性的对称双循环算法。最后,我们将该方法应用于零和马尔可夫博弈,并通过仿真展示了其长期最优性。