A recent paper by Farina & Pipis (2023) established the existence of uncoupled no-linear-swap regret dynamics with polynomial-time iterations in extensive-form games. The equilibrium points reached by these dynamics, known as linear correlated equilibria, are currently the tightest known relaxation of correlated equilibrium that can be learned in polynomial time in any finite extensive-form game. However, their properties remain vastly unexplored, and their computation is onerous. In this paper, we provide several contributions shedding light on the fundamental nature of linear-swap regret. First, we show a connection between linear deviations and a generalization of communication deviations in which the player can make queries to a "mediator" who replies with action recommendations, and, critically, the player is not constrained to match the timing of the game as would be the case for communication deviations. We coin this latter set the untimed communication (UTC) deviations. We show that the UTC deviations coincide precisely with the linear deviations, and therefore that any player minimizing UTC regret also minimizes linear-swap regret. We then leverage this connection to develop state-of-the-art no-regret algorithms for computing linear correlated equilibria, both in theory and in practice. In theory, our algorithms achieve polynomially better per-iteration runtimes; in practice, our algorithms represent the state of the art by several orders of magnitude.
翻译:Farina & Pipis (2023) 的最新论文确立了在扩展形式博弈中具有多项式时间迭代的非耦合无线性交换遗憾动态的存在性。这些动态达到的均衡点被称为线性相关均衡,是目前已知在任何有限扩展形式博弈中可在多项式时间内学习的最紧相关均衡松弛形式。然而,其性质仍远未得到充分探索,且计算过程异常繁重。本文通过多项贡献揭示了线性交换遗憾的基本性质。首先,我们证明了线性偏离与一种通信偏离的泛化之间的关联:在这种泛化中,玩家可向提供行动建议的"中介"进行查询,且关键的是,玩家不受限于匹配通信偏离中的博弈时序。我们将后者称为非定时通信(UTC)偏离。研究表明,UTC偏离与线性偏离完全一致,因此任何最小化UTC遗憾的玩家也同时最小化了线性交换遗憾。随后,我们利用这一关联开发出理论及实践层面均达到最优水平的无遗憾算法,用于计算线性相关均衡。理论上,我们的算法在每轮迭代中实现了多项式级别的运行时间优化;实践中,算法性能以多个数量级优势领先现有最优方案。