We consider Thompson sampling for linear bandit problems with finitely many independent arms, where rewards are sampled from normal distributions that are linearly dependent on unknown parameter vectors and with unknown variance. Specifically, with a Bayesian formulation we consider multivariate normal-gamma priors to represent environment uncertainty for all involved parameters. We show that our chosen sampling prior is a conjugate prior to the reward model and derive a Bayesian regret bound for Thompson sampling under the condition that the 5/2-moment of the variance distribution exist.
翻译:我们考虑具有有限独立臂的线性赌博机问题的汤普森采样,其中奖励来自正态分布,这些分布线性依赖于未知参数向量且方差未知。具体而言,通过贝叶斯框架,我们采用多元正态-伽马先验来表示所有相关参数的环境不确定性。我们证明所选的采样先验是奖励模型的共轭先验,并在方差分布的5/2阶矩存在的条件下,推导出汤普森采样的贝叶斯遗憾界。