We consider the Scale-Free Adversarial Multi-Armed Bandit (MAB) problem with unrestricted feedback delays. In contrast to the standard assumption that all losses are $[0,1]$-bounded, in our setting, losses can fall in a general bounded interval $[-L, L]$, unknown to the agent beforehand. Furthermore, the feedback of each arm pull can experience arbitrary delays. We propose a novel approach named Scale-Free Delayed INF (SFD-INF) for this novel setting, which combines a recent "convex combination trick" together with a novel doubling and skipping technique. We then present two instances of SFD-INF, each with carefully designed delay-adapted learning scales. The first one SFD-TINF uses $\frac 12$-Tsallis entropy regularizer and can achieve $\widetilde{\mathcal O}(\sqrt{K(D+T)}L)$ regret when the losses are non-negative, where $K$ is the number of actions, $T$ is the number of steps, and $D$ is the total feedback delay. This bound nearly matches the $\Omega((\sqrt{KT}+\sqrt{D\log K})L)$ lower-bound when regarding $K$ as a constant independent of $T$. The second one, SFD-LBINF, works for general scale-free losses and achieves a small-loss style adaptive regret bound $\widetilde{\mathcal O}(\sqrt{K\mathbb{E}[\tilde{\mathfrak L}_T^2]}+\sqrt{KDL})$, which falls to the $\widetilde{\mathcal O}(\sqrt{K(D+T)}L)$ regret in the worst case and is thus more general than SFD-TINF despite a more complicated analysis and several extra logarithmic dependencies. Moreover, both instances also outperform the existing algorithms for non-delayed (i.e., $D=0$) scale-free adversarial MAB problems, which can be of independent interest.
翻译:我们研究具有无限制反馈延迟的无尺度对抗性多臂赌博机(MAB)问题。与传统假设损失函数在$[0,1]$区间内不同,本问题中损失可落于未知的广义有界区间$[-L, L]$,且每次拉臂操作的反馈可能遭受任意延迟。针对这一新场景,我们提出名为无尺度延迟INF(SFD-INF)的原创方法,该方法融合了近期提出的"凸组合技巧"与创新的加倍跳过策略。我们进一步给出SFD-INF的两个实例,每个实例均配有精心设计的延迟自适应学习尺度。第一个实例SFD-TINF采用$\frac 12$-Tsallis熵正则化器,在损失非负条件下可实现$\widetilde{\mathcal O}(\sqrt{K(D+T)}L)$遗憾界,其中$K$为动作数、$T$为步数、$D$为总反馈延迟。该结果与下界$\Omega((\sqrt{KT}+\sqrt{D\log K})L)$(将$K$视为与$T$无关的常数)近乎匹配。第二个实例SFD-LBINF适用于通用无尺度损失,具备小损失型自适应遗憾界$\widetilde{\mathcal O}(\sqrt{K\mathbb{E}[\tilde{\mathfrak L}_T^2]}+\sqrt{KDL})$,其最坏情形下退化为$\widetilde{\mathcal O}(\sqrt{K(D+T)}L)$遗憾界,故虽需更复杂的分析与若干额外对数项,但比SFD-TINF更具普适性。此外,两个实例在无延迟(即$D=0$)的无尺度对抗性MAB问题上同样优于现有算法,这本身可作为独立兴趣点。