Multiplayer bandits have recently been extensively studied because of their application to cognitive radio networks. While the literature mostly considers synchronous players, radio networks (e.g. for IoT) tend to have asynchronous devices. This motivates the harder, asynchronous multiplayer bandits problem, which was first tackled with an explore-then-commit (ETC) algorithm (see Dakdouk, 2022), with a regret upper-bound in $\mathcal{O}(T^{\frac{2}{3}})$. Before even considering decentralization, understanding the centralized case was still a challenge as it was unknown whether getting a regret smaller than $\Omega(T^{\frac{2}{3}})$ was possible. We answer positively this question, as a natural extension of UCB exhibits a $\mathcal{O}(\sqrt{T\log(T)})$ minimax regret. More importantly, we introduce Cautious Greedy, a centralized algorithm that yields constant instance-dependent regret if the optimal policy assigns at least one player on each arm (a situation that is proved to occur when arm means are close enough). Otherwise, its regret increases as the sum of $\log(T)$ over some sub-optimality gaps. We provide lower bounds showing that Cautious Greedy is optimal in the data-dependent terms. Therefore, we set up a strong baseline for asynchronous multiplayer bandits and suggest that learning the optimal policy in this problem might be easier than thought, at least with centralization.
翻译:多人博弈机因其在认知无线电网络中的应用而近年来受到广泛研究。尽管文献大多考虑同步玩家,但无线电网络(例如物联网)往往具有异步设备。这促使了更困难的异步多人博弈机问题,该问题首次通过探索-然后-承诺算法(参见Dakdouk,2022)解决,其遗憾上界为$\mathcal{O}(T^{\frac{2}{3}})$。在考虑去中心化之前,理解中心化情况仍是一个挑战,因为未知是否能够获得小于$\Omega(T^{\frac{2}{3}})$的遗憾。我们对此问题给出肯定回答,自然扩展的UCB算法展现出$\mathcal{O}(\sqrt{T\log(T)})$的极小极大遗憾。更重要的是,我们引入谨慎贪婪算法,这是一种中心化算法,若最优策略在每个臂上分配至少一个玩家(当臂均值足够接近时证明该情况发生),则产生常数实例依赖遗憾。否则,其遗憾随某些次优间隙的$\log(T)$之和增加。我们提供的下界表明,谨慎贪婪算法在数据依赖项上是最优的。因此,我们为异步多人博弈机建立了强基线,并表明学习该问题的最优策略可能比想象中更容易,至少在中心化情况下如此。