Confidence bounds are an essential tool for rigorously quantifying the uncertainty of predictions. In this capacity, they can inform the exploration-exploitation trade-off and form a core component in many sequential learning and decision-making algorithms. Tighter confidence bounds give rise to algorithms with better empirical performance and better performance guarantees. In this work, we use martingale tail bounds and finite-dimensional reformulations of infinite-dimensional convex programs to establish new confidence bounds for sequential kernel regression. We prove that our new confidence bounds are always tighter than existing ones in this setting. We apply our confidence bounds to the kernel bandit problem, where future actions depend on the previous history. When our confidence bounds replace existing ones, the KernelUCB (GP-UCB) algorithm has better empirical performance, a matching worst-case performance guarantee and comparable computational cost. Our new confidence bounds can be used as a generic tool to design improved algorithms for other kernelised learning and decision-making problems.
翻译:置信界是严格量化预测不确定性的关键工具。在此能力下,它们能够指导探索-利用权衡,并成为许多序贯学习与决策制定算法的核心组成部分。更紧的置信界能带来具有更优经验性能与更强性能保证的算法。在本工作中,我们利用鞅尾界以及无限维凸规划的有限维重构,为序贯核回归建立了新的置信界。我们证明,在现有框架下,我们提出的新置信界始终比已有置信界更紧。我们将置信界应用于核赌博机问题,其中未来行动依赖于历史记录。当用我们的置信界替代现有置信界时,KernelUCB (GP-UCB) 算法展现出更优的经验性能、匹配的最坏情况性能保证以及相当的计算成本。我们的新置信界可作为通用工具,用于设计其他核化学习与决策制定问题的改进算法。