We investigate learning the equilibria in non-stationary multi-agent systems and address the challenges that differentiate multi-agent learning from single-agent learning. Specifically, we focus on games with bandit feedback, where testing an equilibrium can result in substantial regret even when the gap to be tested is small, and the existence of multiple optimal solutions (equilibria) in stationary games poses extra challenges. To overcome these obstacles, we propose a versatile black-box approach applicable to a broad spectrum of problems, such as general-sum games, potential games, and Markov games, when equipped with appropriate learning and testing oracles for stationary environments. Our algorithms can achieve $\widetilde{O}\left(\Delta^{1/4}T^{3/4}\right)$ regret when the degree of nonstationarity, as measured by total variation $\Delta$, is known, and $\widetilde{O}\left(\Delta^{1/5}T^{4/5}\right)$ regret when $\Delta$ is unknown, where $T$ is the number of rounds. Meanwhile, our algorithm inherits the favorable dependence on number of agents from the oracles. As a side contribution that may be independent of interest, we show how to test for various types of equilibria by a black-box reduction to single-agent learning, which includes Nash equilibria, correlated equilibria, and coarse correlated equilibria.
翻译:我们研究了非平稳多智能体系统中的均衡学习问题,并阐述了区别于单智能体学习的挑战。具体而言,我们聚焦于具有赌博反馈的博弈:当被测试的均衡差距较小时,测试均衡仍可能导致显著遗憾;此外,平稳博弈中存在多个最优解(均衡)也带来额外困难。为克服这些障碍,我们提出了一种普适的黑盒方法,在配备适用于平稳环境的相应学习与测试预言机后,可广泛应用于一般和博弈、势博弈及马尔可夫博弈等问题。当非平稳程度(以总变差Δ衡量)已知时,我们的算法可实现$\widetilde{O}\left(\Delta^{1/4}T^{3/4}\right)$的遗憾界;当Δ未知时,则可实现$\widetilde{O}\left(\Delta^{1/5}T^{4/5}\right)$的遗憾界,其中T为回合数。同时,算法继承了预言机对智能体数量的优良依赖性质。作为一项可能具有独立价值的辅助贡献,我们展示了如何通过黑盒归约至单智能体学习来测试各类均衡,包括纳什均衡、相关均衡和粗相关均衡。