Despite the great interest in the bandit problem, designing efficient algorithms for complex models remains challenging, as there is typically no analytical way to quantify uncertainty. In this paper, we propose Multiplier Bootstrap-based Exploration (MBE), a novel exploration strategy that is applicable to any reward model amenable to weighted loss minimization. We prove both instance-dependent and instance-independent rate-optimal regret bounds for MBE in sub-Gaussian multi-armed bandits. With extensive simulation and real data experiments, we show the generality and adaptivity of MBE.
翻译:尽管对赌博机问题有着极大的研究兴趣,但为复杂模型设计高效算法仍具挑战性,因为通常不存在解析方法来量化不确定性。本文提出名为乘子自助法探索(MBE)的新型探索策略,该策略适用于任何可进行加权损失最小化的奖励模型。我们证明了在次高斯多臂赌博机中MBE具有实例相关和实例无关的率最优遗憾界。通过大量模拟实验和真实数据实验,我们展示了MBE的通用性与自适应性。