The study of collaborative multi-agent bandits has attracted significant attention recently. In light of this, we initiate the study of a new collaborative setting, consisting of $N$ agents such that each agent is learning one of $M$ stochastic multi-armed bandits to minimize their group cumulative regret. We develop decentralized algorithms which facilitate collaboration between the agents under two scenarios. We characterize the performance of these algorithms by deriving the per agent cumulative regret and group regret upper bounds. We also prove lower bounds for the group regret in this setting, which demonstrates the near-optimal behavior of the proposed algorithms.
翻译:协作式多智能体赌博机的研究近期引起了广泛关注。为此,我们率先研究了一种新的协作设置,该设置包含N个智能体,每个智能体分别学习M个随机多臂赌博机之一,以最小化其群体累积遗憾。我们开发了两种场景下的去中心化算法,以促进智能体间的协作。通过推导个体智能体累积遗憾和群体遗憾的上界,我们刻画了这些算法的性能。我们还证明了该设置下群体遗憾的下界,这表明所提出算法具有近乎最优的性能。