Thompson sampling has proven effective across a wide range of stationary bandit environments. However, as we demonstrate in this paper, it can perform poorly when applied to non-stationary environments. We show that such failures are attributed to the fact that, when exploring, the algorithm does not differentiate actions based on how quickly the information acquired loses its usefulness due to nonstationarity. Building upon this insight, we propose predictive sampling, an algorithm that deprioritizes acquiring information that quickly loses usefulness. Theoretical guarantee on the performance of predictive sampling is established through a Bayesian regret bound. We provide versions of predictive sampling for which computations tractably scale to complex bandit environments of practical interest. Through numerical simulations, we demonstrate that predictive sampling outperforms Thompson sampling in all non-stationary environments examined.
翻译:汤普森采样已被证明在广泛平稳赌博机环境中行之有效。然而,正如本文所示,当应用于非平稳环境时,其性能可能非常糟糕。我们发现这种失败归因于算法在探索时,未能根据信息因非平稳性而丧失效用的速度来区分不同动作。基于这一见解,我们提出预测采样算法,该算法降低了对快速失效信息的获取优先级。通过贝叶斯遗憾界建立了预测采样性能的理论保证。我们提供了适用于实际复杂赌博机环境且计算可扩展的预测采样版本。数值仿真表明,在所有被考察的非平稳环境中,预测采样表现均优于汤普森采样。