For the adversarial multi-armed bandit problem with delayed feedback, we consider that the delayed feedback results are from multiple users and are unrestricted on internal distribution. As the player picks an arm, feedback from multiple users may not be received instantly yet after an arbitrary delay of time which is unknown to the player in advance. For different users in a round, the delays in feedback have no latent correlation. Thus, we formulate an adversarial multi-armed bandit problem with multi-user delayed feedback and design a modified EXP3 algorithm named MUD-EXP3, which makes a decision at each round by considering the importance-weighted estimator of the received feedback from different users. On the premise of known terminal round index $T$, the number of users $M$, the number of arms $N$, and upper bound of delay $d_{max}$, we prove a regret of $\mathcal{O}(\sqrt{TM^2\ln{N}(N\mathrm{e}+4d_{max})})$. Furthermore, for the more common case of unknown $T$, an adaptive algorithm named AMUD-EXP3 is proposed with a sublinear regret with respect to $T$. Finally, extensive experiments are conducted to indicate the correctness and effectiveness of our algorithms.
翻译:摘要:针对具有延迟反馈的对抗性多臂赌博机问题,我们考虑延迟反馈结果来自多个用户且内部分布不受限制的情形。当玩家选择一根臂时,来自多个用户的反馈不会立即收到,而是经历一个玩家事先未知的任意延迟时间。同一轮次中不同用户的反馈延迟不存在潜在相关性。由此,我们构建了具有多用户延迟反馈的对抗性多臂赌博机问题,并设计了一种名为MUD-EXP3的改进型EXP3算法,该算法通过考虑来自不同用户的已接收反馈的重要性加权估计量,在每个轮次做出决策。在已知终止轮次索引$T$、用户数量$M$、臂数$N$以及延迟上界$d_{max}$的前提下,我们证明了其遗憾值为$\mathcal{O}(\sqrt{TM^2\ln{N}(N\mathrm{e}+4d_{max})})$。进一步地,针对更常见的$T$未知情形,提出了一种自适应算法AMUD-EXP3,其遗憾值关于$T$呈次线性增长。最后,通过大量实验验证了所提算法的正确性与有效性。