We provide the first useful, rigorous analysis of ensemble sampling for the stochastic linear bandit setting. In particular, we show that, under standard assumptions, for a $d$-dimensional stochastic linear bandit with an interaction horizon $T$, ensemble sampling with an ensemble of size $m$ on the order of $d \log T$ incurs regret bounded by order $(d \log T)^{5/2} \sqrt{T}$. Ours is the first result in any structured setting not to require the size of the ensemble to scale linearly with $T$ -- which defeats the purpose of ensemble sampling -- while obtaining near $\sqrt{T}$ order regret. Ours is also the first result that allows infinite action sets.
翻译:我们首次对随机线性赌博机场景下的集成采样进行了实用且严谨的分析。具体而言,我们证明在标准假设下,对于一个交互时长为$T$的$d$维随机线性赌博机,当集成规模$m$量级为$d \log T$时,集成采样所导致的遗憾上限为$(d \log T)^{5/2} \sqrt{T}$。这是任何结构化场景中首个无需集成规模与$T$成线性关系(否则将违背集成采样的初衷)且能达到接近$\sqrt{T}$量级遗憾的结果。此外,本研究也是首个允许无限动作集的结果。