A unique challenge in Multi-Agent Reinforcement Learning (MARL) is the curse of multiagency, where the description length of the game as well as the complexity of many existing learning algorithms scale exponentially with the number of agents. While recent works successfully address this challenge under the model of tabular Markov Games, their mechanisms critically rely on the number of states being finite and small, and do not extend to practical scenarios with enormous state spaces where function approximation must be used to approximate value functions or policies. This paper presents the first line of MARL algorithms that provably resolve the curse of multiagency under function approximation. We design a new decentralized algorithm -- V-Learning with Policy Replay, which gives the first polynomial sample complexity results for learning approximate Coarse Correlated Equilibria (CCEs) of Markov Games under decentralized linear function approximation. Our algorithm always outputs Markov CCEs, and achieves an optimal rate of $\widetilde{\mathcal{O}}(\epsilon^{-2})$ for finding $\epsilon$-optimal solutions. Also, when restricted to the tabular case, our result improves over the current best decentralized result $\widetilde{\mathcal{O}}(\epsilon^{-3})$ for finding Markov CCEs. We further present an alternative algorithm -- Decentralized Optimistic Policy Mirror Descent, which finds policy-class-restricted CCEs using a polynomial number of samples. In exchange for learning a weaker version of CCEs, this algorithm applies to a wider range of problems under generic function approximation, such as linear quadratic games and MARL problems with low ''marginal'' Eluder dimension.
翻译:多智能体强化学习(MARL)中的一个独特挑战是多智能体诅咒,即博弈的描述长度以及许多现有学习算法的复杂度会随智能体数量呈指数级增长。尽管近期研究在表格型马尔可夫博弈模型下成功解决了这一挑战,但其机制关键依赖于状态空间有限且较小,无法扩展至需要借助函数逼近来近似值函数或策略的、具有庞大状态空间的实际场景。本文提出了首个在函数逼近下可证明解决多智能体诅咒的MARL算法系列。我们设计了一种新的去中心化算法——带策略回放的V学习(V-Learning with Policy Replay),该算法在去中心化线性函数逼近下,首次获得了学习马尔可夫博弈近似粗糙相关均衡(CCEs)的多项式样本复杂度结果。我们的算法始终输出马尔可夫CCE,并以最优速率$\widetilde{\mathcal{O}}(\epsilon^{-2})$找到$\epsilon$-最优解。此外,在表格型场景下,我们的结果将寻找马尔可夫CCE的当前最优去中心化结果从$\widetilde{\mathcal{O}}(\epsilon^{-3})$改进至$\widetilde{\mathcal{O}}(\epsilon^{-2})$。我们进一步提出另一种替代算法——去中心化乐观策略镜像下降(Decentralized Optimistic Policy Mirror Descent),该算法使用多项式数量的样本找到策略类受限的CCE。作为学习较弱版本CCE的代价,此算法适用于更广泛的问题,包括线性二次博弈和具有低“边际”艾尔维度(Eluder dimension)的MARL问题。