We study a distributed stochastic multi-armed bandit where a client supplies the learner with communication-constrained feedback based on the rewards for the corresponding arm pulls. In our setup, the client must encode the rewards such that the second moment of the encoded rewards is no more than $P$, and this encoded reward is further corrupted by additive Gaussian noise of variance $\sigma^2$; the learner only has access to this corrupted reward. For this setting, we derive an information-theoretic lower bound of $\Omega\left(\sqrt{\frac{KT}{\mathtt{SNR} \wedge1}} \right)$ on the minimax regret of any scheme, where $ \mathtt{SNR} := \frac{P}{\sigma^2}$, and $K$ and $T$ are the number of arms and time horizon, respectively. Furthermore, we propose a multi-phase bandit algorithm, $\mathtt{UE\text{-}UCB++}$, which matches this lower bound to a minor additive factor. $\mathtt{UE\text{-}UCB++}$ performs uniform exploration in its initial phases and then utilizes the {\em upper confidence bound }(UCB) bandit algorithm in its final phase. An interesting feature of $\mathtt{UE\text{-}UCB++}$ is that the coarser estimates of the mean rewards formed during a uniform exploration phase help to refine the encoding protocol in the next phase, leading to more accurate mean estimates of the rewards in the subsequent phase. This positive reinforcement cycle is critical to reducing the number of uniform exploration rounds and closely matching our lower bound.
翻译:我们研究了一种分布式随机多臂土匪问题,其中客户端基于对应臂抽奖获得的奖励向学习器提供通信受限的反馈。在我们的设置中,客户端必须对奖励进行编码,使得编码奖励的二阶矩不超过 $P$,且该编码奖励进一步被方差为 $\sigma^2$ 的加性高斯噪声所污染;学习器仅能访问该污染后的奖励。针对该场景,我们推导出任何方案的最小最大遗憾的信息理论下界为 $\Omega\left(\sqrt{\frac{KT}{\mathtt{SNR} \wedge1}} \right)$,其中 $\mathtt{SNR} := \frac{P}{\sigma^2}$,$K$ 和 $T$ 分别表示臂数与时间范围。此外,我们提出了一种多阶段强盗算法 $\mathtt{UE\text{-}UCB++}$,该算法与下界仅相差一个较小的加性因子。$\mathtt{UE\text{-}UCB++}$ 在初始阶段执行均匀探索,并在最终阶段采用上置信界(UCB)强盗算法。该算法的一个有趣特性是:均匀探索阶段形成的均值奖励粗估计有助于优化下一阶段的编码协议,从而在后续阶段获得更精准的均值奖励估计。这种正反馈循环对减少均匀探索轮次并紧密贴合下界至关重要。