We consider a decentralized multiplayer game, played over $T$ rounds, with a leader-follower hierarchy described by a directed acyclic graph. For each round, the graph structure dictates the order of the players and how players observe the actions of one another. By the end of each round, all players receive a joint bandit-reward based on their joint action that is used to update the player strategies towards the goal of minimizing the joint pseudo-regret. We present a learning algorithm inspired by the single-player multi-armed bandit problem and show that it achieves sub-linear joint pseudo-regret in the number of rounds for both adversarial and stochastic bandit rewards. Furthermore, we quantify the cost incurred due to the decentralized nature of our problem compared to the centralized setting.
翻译:我们考虑一个在$T$轮博弈中进行的分散式多人博弈,其领导者-跟随者层级由有向无环图描述。每轮博弈中,图结构决定了玩家的行动顺序以及玩家之间如何观察彼此的行动。每轮结束时,所有玩家根据其联合行动获得一个联合赌博机奖励,该奖励用于更新玩家策略,以最小化联合伪遗憾为目标。我们提出了一种受单玩家多臂赌博机问题启发的学习算法,并证明该算法在面对对抗性奖励和随机性奖励时,均能实现关于轮数的次线性联合伪遗憾。此外,我们量化了相较于集中式场景,因问题的分散式特性所导致的额外代价。