This paper studies a cooperative multi-agent multi-armed stochastic bandit problem where agents operate asynchronously -- agent pull times and rates are unknown, irregular, and heterogeneous -- and face the same instance of a K-armed bandit problem. Agents can share reward information to speed up the learning process at additional communication costs. We propose ODC, an on-demand communication protocol that tailors the communication of each pair of agents based on their empirical pull times. ODC is efficient when the pull times of agents are highly heterogeneous, and its communication complexity depends on the empirical pull times of agents. ODC is a generic protocol that can be integrated into most cooperative bandit algorithms without degrading their performance. We then incorporate ODC into the natural extensions of UCB and AAE algorithms and propose two communication-efficient cooperative algorithms. Our analysis shows that both algorithms are near-optimal in regret.
翻译:本文研究一个合作式多智能体多臂随机赌博机问题,其中智能体异步运行——各智能体的拉臂时间和速率未知、不规则且异质——并面对同一个K臂赌博机实例。智能体可通过共享奖励信息来加速学习过程,但需承担额外通信成本。我们提出ODC(按需通信协议),该协议根据每对智能体的实际拉臂时间为其定制通信策略。当智能体拉臂时间高度异质时,ODC具有高效性,且其通信复杂度取决于智能体的实际拉臂时间。ODC是一种通用协议,可集成至多数合作式赌博机算法中而不降低其性能。我们随后将ODC融入UCB和AAE算法的自然扩展版本,提出两种通信高效的协同算法。理论分析表明,两种算法的遗憾值均接近最优。