We consider the stochastic linear contextual bandit problem with high-dimensional features. We analyze the Thompson sampling algorithm using special classes of sparsity-inducing priors (e.g., spike-and-slab) to model the unknown parameter and provide a nearly optimal upper bound on the expected cumulative regret. To the best of our knowledge, this is the first work that provides theoretical guarantees of Thompson sampling in high-dimensional and sparse contextual bandits. For faster computation, we use variational inference instead of Markov Chain Monte Carlo (MCMC) to approximate the posterior distribution. Extensive simulations demonstrate the improved performance of our proposed algorithm over existing ones.
翻译:我们研究了高维特征下的随机线性上下文赌博机问题。通过采用特定类型的稀疏诱导先验(例如,spike-and-slab先验)对未知参数进行建模,我们分析了汤普森采样算法,并给出了期望累积遗憾的近乎最优上界。据我们所知,这是首个为高维稀疏上下文赌博机中汤普森采样提供理论保证的工作。为了提升计算效率,我们采用变分推断代替马尔可夫链蒙特卡洛(MCMC)方法来近似后验分布。广泛模拟表明,我们提出的算法相较于现有算法具有更优性能。