We study the performance of the gradient play algorithm for stochastic games (SGs), where each agent tries to maximize its own total discounted reward by making decisions independently based on current state information which is shared between agents. Policies are directly parameterized by the probability of choosing a certain action at a given state. We show that Nash equilibria (NEs) and first-order stationary policies are equivalent in this setting, and give a local convergence rate around strict NEs. Further, for a subclass of SGs called Markov potential games (which includes the setting with identical rewards as an important special case), we design a sample-based reinforcement learning algorithm and give a non-asymptotic global convergence rate analysis for both exact gradient play and our sample-based learning algorithm. Our result shows that the number of iterations to reach an $\epsilon$-NE scales linearly, instead of exponentially, with the number of agents. Local geometry and local stability are also considered, where we prove that strict NEs are local maxima of the total potential function and fully-mixed NEs are saddle points.
翻译:我们研究了随机博弈中梯度博弈算法的性能,其中每个智能体基于共享的当前状态信息独立决策,以最大化自身总折现奖励。策略通过给定状态下选择特定动作的概率直接参数化。我们证明在该设定下纳什均衡与一阶驻点是等价的,并给出了严格纳什均衡附近的局部收敛速率。进一步,针对一类称为马尔可夫势博弈的随机博弈子类(包含具有相同奖励的重要特例),我们设计了一种基于样本的强化学习算法,并对精确梯度博弈及基于样本的学习算法给出了非渐近全局收敛速率分析。结果表明,达到$\epsilon$-纳什均衡所需的迭代次数随智能体数量呈线性增长(而非指数增长)。我们还考虑了局部几何与局部稳定性,证明了严格纳什均衡是总势函数的局部极大值点,而完全混合纳什均衡是鞍点。