This work focuses on the entropy-regularized independent natural policy gradient (NPG) algorithm in multi-agent reinforcement learning. In this work, agents are assumed to have access to an oracle with exact policy evaluation and seek to maximize their respective independent rewards. Each individual's reward is assumed to depend on the actions of all the agents in the multi-agent system, leading to a game between agents. We assume all agents make decisions under a policy with bounded rationality, which is enforced by the introduction of entropy regularization. In practice, a smaller regularization implies the agents are more rational and behave closer to Nash policies. On the other hand, agents with larger regularization acts more randomly, which ensures more exploration. We show that, under sufficient entropy regularization, the dynamics of this system converge at a linear rate to the quantal response equilibrium (QRE). Although regularization assumptions prevent the QRE from approximating a Nash equilibrium, our findings apply to a wide range of games, including cooperative, potential, and two-player matrix games. We also provide extensive empirical results on multiple games (including Markov games) as a verification of our theoretical analysis.
翻译:本文聚焦于多智能体强化学习中的熵正则化独立自然策略梯度(NPG)算法。假设智能体可访问精确策略评估的预言机,并致力于最大化各自的独立奖励。每个智能体的奖励依赖于多智能体系统中所有智能体的行为,从而形成智能体间的博弈。我们假设所有智能体在有限理性约束下做出决策,该约束通过引入熵正则化实现。实际中,较小的正则化意味着智能体更理性,行为更接近纳什策略;而较大的正则化则使智能体行为更具随机性,从而保证更充分的探索。研究表明,在充分熵正则化条件下,该系统的动态过程以线性速率收敛至分位响应均衡(QRE)。尽管正则化假设阻止了QRE逼近纳什均衡,但本文结论适用于包括合作博弈、势博弈及双人矩阵博弈在内的广泛博弈类型。我们还提供了多项博弈(包括马尔可夫博弈)的丰富实证结果,以验证理论分析的正确性。