The recent literature on online learning to rank (LTR) has established the utility of prior knowledge to Bayesian ranking bandit algorithms. However, a major limitation of existing work is the requirement for the prior used by the algorithm to match the true prior. In this paper, we propose and analyze adaptive algorithms that address this issue and additionally extend these results to the linear and generalized linear models. We also consider scalar relevance feedback on top of click feedback. Moreover, we demonstrate the efficacy of our algorithms using both synthetic and real-world experiments.
翻译:近期关于在线排序学习(LTR)的文献已证实先验知识对贝叶斯排序赌博机算法的效用。然而,现有研究的一个主要局限是要求算法使用的先验分布必须与真实先验分布相匹配。本文提出并分析了一类自适应算法以解决该问题,并进一步将这些结果推广至线性模型与广义线性模型。我们还考虑了在点击反馈基础上引入标量相关性反馈。此外,通过合成数据实验和真实场景实验验证了所提算法的有效性。