We study nonparametric contextual bandits where Lipschitz mean reward functions may change over time. We first establish the minimax dynamic regret rate in this less understood setting in terms of number of changes $L$ and total-variation $V$, both capturing all changes in distribution over context space, and argue that state-of-the-art procedures are suboptimal in this setting. Next, we tend to the question of an adaptivity for this setting, i.e. achieving the minimax rate without knowledge of $L$ or $V$. Quite importantly, we posit that the bandit problem, viewed locally at a given context $X_t$, should not be affected by reward changes in other parts of context space $\cal X$. We therefore propose a notion of change, which we term experienced significant shifts, that better accounts for locality, and thus counts considerably less changes than $L$ and $V$. Furthermore, similar to recent work on non-stationary MAB (Suk & Kpotufe, 2022), experienced significant shifts only count the most significant changes in mean rewards, e.g., severe best-arm changes relevant to observed contexts. Our main result is to show that this more tolerant notion of change can in fact be adapted to.
翻译:我们研究非参数化情境赌博机问题,其中Lipschitz均值奖励函数可能随时间变化。首先,在这一尚未充分理解的场景下,我们建立了关于变化次数$L$和总变差$V$(两者均捕捉情境空间上分布的所有变化)的极小化动态遗憾率,并论证了现有先进算法在该场景下的次优性。接着,我们探讨该场景的自适应性,即无需知晓$L$或$V$即可达到极小化率。重要的是,我们提出:在给定情境$X_t$处局部看待的赌博机问题,不应受到情境空间$\cal X$其他部分奖励变化的影响。因此,我们提出一种称为"经验显著变迁"的变化概念,它能更好地考虑局部性,从而统计出远少于$L$和$V$的变化次数。此外,类似于非平稳MAB的最新工作(Suk & Kpotufe, 2022),经验显著变迁仅统计均值奖励中最显著的变化,例如与观测情境相关的严重最优臂变化。我们的主要结果是证明这种更具宽容性的变化概念实际上是可以自适应处理的。