In this paper, we study the problem of fair multi-agent multi-arm bandit learning when agents do not communicate with each other, except collision information, provided to agents accessing the same arm simultaneously. We provide an algorithm with regret $O\left(N^3 \log \frac{B}{\Delta} f(\log T) \log T \right)$ (assuming bounded rewards, with unknown bound), where $f(t)$ is any function diverging to infinity with $t$. This significantly improves previous results which had the same upper bound on the regret of order $O(f(\log T) \log T )$ but an exponential dependence on the number of agents. The result is attained by using a distributed auction algorithm to learn the sample-optimal matching and a novel order-statistics-based regret analysis. Simulation results present the dependence of the regret on $\log T$.
翻译:在本文中,我们研究了当智能体之间不进行通信(除同时访问同一臂时提供的碰撞信息外)的公平多智能体多臂赌博机学习问题。我们提出了一种算法,其遗憾界为 $O\left(N^3 \log \frac{B}{\Delta} f(\log T) \log T \right)$(假设奖励有界,且界未知),其中 $f(t)$ 是任意随 $t$ 趋于无穷大的函数。这显著改进了先前的结果,后者虽具有相同的上界 $O(f(\log T) \log T )$,但对智能体数量呈指数依赖。该结果通过采用分布式拍卖算法学习样本最优匹配,并基于顺序统计量的新型遗憾分析实现。仿真结果展示了遗憾对 $\log T$ 的依赖关系。