This paper addresses the problem of learning Nash equilibria in {\it monotone games} where the gradient of the payoff functions is monotone in the strategy profile space, potentially containing additive noise. The optimistic family of learning algorithms, exemplified by optimistic Follow-the-Regularized-Leader and optimistic Mirror Descent, successfully achieves last-iterate convergence in scenarios devoid of noise, leading the dynamics to a Nash equilibrium. A recent emerging trend underscores the promise of the perturbation approach, where payoff functions are perturbed based on the distance from an anchoring, or {\it slingshot}, strategy. In response, we first establish a unified framework for learning equilibria in monotone games, accommodating both full and noisy feedback. Second, we construct the convergence rates toward an approximated equilibrium, irrespective of noise presence. Thirdly, we introduce a twist by updating the slingshot strategy, anchoring the current strategy at finite intervals. This innovation empowers us to identify the exact Nash equilibrium of the underlying game with guaranteed rates. The proposed framework is all-encompassing, integrating existing payoff-perturbed algorithms. Finally, empirical demonstrations affirm that our algorithms, grounded in this framework, exhibit significantly accelerated convergence.
翻译:本文研究了在{\it单调博弈}中学习纳什均衡的问题,其中支付函数的梯度在策略配置空间上是单调的,且可能包含加性噪声。以乐观跟随正则化领导者和乐观镜像下降为代表的乐观学习算法,在无噪声场景下成功实现了末次迭代收敛,使动力学系统趋向纳什均衡。最近的新兴趋势凸显了扰动方法的潜力,该方法根据与锚定(或{\it弹弓})策略的距离对支付函数进行扰动。为此,我们首先建立了一个统一框架用于单调博弈中的均衡学习,该框架同时适用于完全反馈和含噪反馈。其次,我们构建了趋近近似均衡的收敛速率,无论噪声是否存在。第三,我们引入一项创新:通过更新弹弓策略,在有限区间内锚定当前策略。这一创新使我们能够以有保证的速率识别基础博弈的精确纳什均衡。所提出的框架具有全面性,整合了现有的支付扰动算法。最后,实证演示表明,基于该框架的算法展现出显著加速的收敛性能。