Competitive multi-agent reinforcement learning in imperfect-information games requires agents to act under partial observability and against adversarial opponents, necessitating stochastic policies. While self-play reinforcement learning with Proximal Policy Optimization (PPO) has achieved strong empirical success, its standard advantage estimator, generalized advantage estimation, suffers from additional variance due to the sampling of stochastic future actions. This variance is amplified in equilibrium self-play because of the stochastic nature of the equilibrium policy and persists even when the critic is exact. We address this bottleneck by introducing $Q$-boosting, a variance-reduced advantage estimator based on a centralized action-value critic, and propose Variance-Reduced Policy Optimization (VRPO), incorporating this new estimator. The algorithm replaces sampled multi-step backups with a multi-step Expected SARSA$(λ)$ trace, computing policy expectations at each step to average out action-sampling noise, while retaining PPO's clipped objective and on-policy actor updates. Empirically, VRPO consistently achieves strong performance from mid-sized to large-scale games including Dou Dizhu and Heads-Up No-Limit Texas Hold'em.
翻译:竞争性多智能体强化学习在不完美信息游戏中要求智能体在部分可观测性和对抗性对手条件下行动,因此必须采用随机策略。尽管基于近端策略优化的自我对弈强化学习已取得显著实证成功,但其标准优势估计器——广义优势估计——因采样随机未来动作而引入额外方差。这一方差在均衡自我对弈过程中因均衡策略的随机性而放大,即使评论家完全精确时依然存在。我们通过引入基于集中式动作-价值评论家的方差缩减优势估计器$Q$-boosting解决这一瓶颈,并提出了集成新估计器的方差缩减策略优化算法。该算法将采样的多步备份替换为多步期望SARSA$(λ)$迹,在每一步计算策略期望以平均消除动作采样噪声,同时保留PPO的裁剪目标和在策略演员更新机制。实验表明,VRPO在包括斗地主和一对一无限注德州扑克的中型至大规模游戏中持续取得优异表现。