We introduce an online learning algorithm in the bandit feedback model that, once adopted by all agents of a congestion game, results in game-dynamics that converge to an $\epsilon$-approximate Nash Equilibrium in a polynomial number of rounds with respect to $1/\epsilon$, the number of players and the number of available resources. The proposed algorithm also guarantees sublinear regret to any agent adopting it. As a result, our work answers an open question from arXiv:2206.01880 and extends the recent results of arXiv:2306.15543 to the bandit feedback model. We additionally establish that our online learning algorithm can be implemented in polynomial time for the important special case of Network Congestion Games on Directed Acyclic Graphs (DAG) by constructing an exact $1$-barycentric spanner for DAGs.
翻译:我们提出了一种在赌博机反馈模型下的在线学习算法。一旦拥塞博弈中的所有智能体采用该算法,博弈动力学将在多项式轮数内收敛到ϵ-近似纳什均衡,其中轮数关于1/ϵ、玩家数量及可用资源数量均为多项式。所提算法还能确保任何采用它的智能体获得次线性遗憾。因此,我们的工作回答了arXiv:2206.01880中的一个开放性问题,并将arXiv:2306.15543的最新成果扩展到了赌博机反馈模型。此外,我们进一步证明,通过为有向无环图构建精确的1-重心张量,该在线学习算法可以在有向无环图网络拥塞博弈这一重要特例中以多项式时间实现。