A unique challenge in Multi-Agent Reinforcement Learning (MARL) is the curse of multiagency, where the description length of the game as well as the complexity of many existing learning algorithms scale exponentially with the number of agents. While recent works successfully address this challenge under the model of tabular Markov Games, their mechanisms critically rely on the number of states being finite and small, and do not extend to practical scenarios with enormous state spaces where function approximation must be used to approximate value functions or policies. This paper presents the first line of MARL algorithms that provably resolve the curse of multiagency under function approximation. We design a new decentralized algorithm -- V-Learning with Policy Replay, which gives the first polynomial sample complexity results for learning approximate Coarse Correlated Equilibria (CCEs) of Markov Games under decentralized linear function approximation. Our algorithm always outputs Markov CCEs, and achieves an optimal rate of $\widetilde{\mathcal{O}}(\epsilon^{-2})$ for finding $\epsilon$-optimal solutions. Also, when restricted to the tabular case, our result improves over the current best decentralized result $\widetilde{\mathcal{O}}(\epsilon^{-3})$ for finding Markov CCEs. We further present an alternative algorithm -- Decentralized Optimistic Policy Mirror Descent, which finds policy-class-restricted CCEs using a polynomial number of samples. In exchange for learning a weaker version of CCEs, this algorithm applies to a wider range of problems under generic function approximation, such as linear quadratic games and MARL problems with low ''marginal'' Eluder dimension.
翻译:多智能体强化学习(MARL)面临的一个独特挑战是“多智能体诅咒”——游戏的描述长度以及许多现有学习算法的复杂度会随智能体数量呈指数级增长。尽管近期研究在表格型马尔可夫博弈模型下成功解决了这一挑战,但这些方法的关键依赖是状态空间有限且较小,无法扩展到需要依赖函数逼近来近似值函数或策略的巨大状态空间实际场景。本文提出了首个在函数逼近下可证明解决多智能体诅咒的MARL算法体系。我们设计了一种新的去中心化算法——基于策略回放的V-Learning,该算法在去中心化线性函数逼近下,首次给出了学习马尔可夫博弈近似粗相关均衡(CCEs)的多项式样本复杂度结果。我们的算法始终输出马尔可夫CCEs,并以最优速率$\widetilde{\mathcal{O}}(\epsilon^{-2})$找到$\epsilon$-最优解。此外,在表格型情形下,我们的结果将寻找马尔可夫CCEs的当前最优去中心化结果从$\widetilde{\mathcal{O}}(\epsilon^{-3})$提升至$\widetilde{\mathcal{O}}(\epsilon^{-2})$。我们进一步提出另一种算法——去中心化乐观策略镜像梯度下降,该算法使用多项式数量的样本即可找到策略类受限的CCEs。作为交换,这种算法虽学习的是较弱的CCEs版本,但适用于更广泛的问题,包括线性二次博弈和具有低“边际”Eluder维度的MARL问题。