We study the generalized linear contextual bandit problem within the requirements of limited adaptivity. In this paper, we present two algorithms, B-GLinCB and RS-GLinCB, that address, respectively, two prevalent limited adaptivity models: batch learning with stochastic contexts and rare policy switches with adversarial contexts. For both these models, we establish essentially tight regret bounds. Notably, in the obtained bounds, we manage to eliminate a dependence on a key parameter $\kappa$, which captures the non-linearity of the underlying reward model. For our batch learning algorithm B-GLinCB, with $\Omega\left( \log{\log T} \right)$ batches, the regret scales as $\tilde{O}(\sqrt{T})$. Further, we establish that our rarely switching algorithm RS-GLinCB updates its policy at most $\tilde{O}(\log^2 T)$ times and achieves a regret of $\tilde{O}(\sqrt{T})$. Our approach for removing the dependence on $\kappa$ for generalized linear contextual bandits might be of independent interest.
翻译:我们研究了在有限适应性要求下的广义线性上下文强盗问题。本文提出了两种算法,B-GLinCB和RS-GLinCB,分别针对两种常见的有限适应性模型:随机上下文下的批量学习和对立上下文下的罕见策略切换。对于这两种模型,我们建立了基本紧致的遗憾界。特别地,在所得界中,我们成功消除了对关键参数$\kappa$的依赖,该参数捕捉了底层奖励模型的非线性。对于我们的批量学习算法B-GLinCB,在$\Omega\left( \log{\log T} \right)$个批量下,遗憾量级为$\tilde{O}(\sqrt{T})$。此外,我们证明了罕见切换算法RS-GLinCB最多更新策略$\tilde{O}(\log^2 T)$次,并实现了$\tilde{O}(\sqrt{T})$的遗憾。我们消除广义线性上下文强盗问题中$\kappa$依赖性的方法可能具有独立的研究价值。