In this work, we study the performance of the Thompson Sampling algorithm for Contextual Bandit problems based on the framework introduced by Neu et al. and their concept of lifted information ratio. First, we prove a comprehensive bound on the Thompson Sampling expected cumulative regret that depends on the mutual information of the environment parameters and the history. Then, we introduce new bounds on the lifted information ratio that hold for sub-Gaussian rewards, thus generalizing the results from Neu et al. which analysis requires binary rewards. Finally, we provide explicit regret bounds for the special cases of unstructured bounded contextual bandits, structured bounded contextual bandits with Laplace likelihood, structured Bernoulli bandits, and bounded linear contextual bandits.
翻译:本文中,我们基于Neu等人提出的框架及其提升信息比概念,研究了汤普森采样算法在上下文赌博机问题中的性能。首先,我们证明了汤普森采样期望累积遗憾的一个全面界,该界依赖于环境参数与历史信息的互信息。随后,我们引入了适用于sub-Gaussian奖励的提升信息比的新界,从而推广了Neu等人仅需二元奖励分析的结果。最后,我们针对无结构有界上下文赌博机、具有拉普拉斯似然的结构化有界上下文赌博机、结构化伯努利赌博机以及有界线性上下文赌博机等特例,给出了明确的遗憾界。