Bandit learning has been an increasingly popular design choice for recommender system. Despite the strong interest in bandit learning from the community, there remains multiple bottlenecks that prevent many bandit learning approaches from productionalization. One major bottleneck is how to test the effectiveness of bandit algorithm with fairness and without data leakage. Different from supervised learning algorithms, bandit learning algorithms emphasize greatly on the data collection process through their explorative nature. Such explorative behavior may induce unfair evaluation in a classic A/B test setting. In this work, we apply upper confidence bound (UCB) to our large scale short video recommender system and present a test framework for the production bandit learning life-cycle with a new set of metrics. Extensive experiment results show that our experiment design is able to fairly evaluate the performance of bandit learning in the recommender system.
翻译:波段学习已成为推荐系统中日益流行的设计选择。尽管学界对波段学习兴趣浓厚,但仍存在多重瓶颈阻碍许多波段学习方法实现产品化。其中一个主要瓶颈是如何公平且无数据泄露地测试波段算法的有效性。与监督学习算法不同,波段学习算法因其探索性特征而高度强调数据收集过程。这种探索性行为可能在经典A/B测试设置中引发不公平评估。在本工作中,我们将上置信界(UCB)应用于大规模短视频推荐系统,并提出一个面向产品级波段学习生命周期的测试框架,配套全新评估指标集。大量实验结果表明,我们的实验设计能够公平地评估推荐系统中波段学习的性能。