In this paper we propose a general methodology to derive regret bounds for randomized multi-armed bandit algorithms. It consists in checking a set of sufficient conditions on the sampling probability of each arm and on the family of distributions to prove a logarithmic regret. As a direct application we revisit two famous bandit algorithms, Minimum Empirical Divergence (MED) and Thompson Sampling (TS), under various models for the distributions including single parameter exponential families, Gaussian distributions, bounded distributions, or distributions satisfying some conditions on their moments. In particular, we prove that MED is asymptotically optimal for all these models, but also provide a simple regret analysis of some TS algorithms for which the optimality is already known. We then further illustrate the interest of our approach, by analyzing a new Non-Parametric TS algorithm (h-NPTS), adapted to some families of unbounded reward distributions with a bounded h-moment. This model can for instance capture some non-parametric families of distributions whose variance is upper bounded by a known constant.
翻译:本文提出了一种通用方法,用于推导随机化多臂老虎机算法的遗憾界。该方法通过检查每个臂的采样概率及分布族的一组充分条件,证明了算法具有对数遗憾。作为直接应用,本文重新审视了两种经典老虎机算法——最小经验散度(MED)和汤普森采样(TS),并考虑了多种分布模型,包括单参数指数族、高斯分布、有界分布以及满足矩条件的分布。特别地,我们证明了MED在这些模型下渐近最优,同时也为部分已知最优性的TS算法提供了简单的遗憾分析。进一步地,我们通过分析一种新的非参数TS算法(h-NPTS)来展示我们方法的优势,该算法适用于具有有界h矩的无界奖励分布族。这类模型可以捕捉一些方差受已知常数上界约束的非参数分布族。