We study the ranking problem in generalized linear bandits. At each time, the learning agent selects an ordered list of items and observes stochastic outcomes. In recommendation systems, displaying an ordered list of the most attractive items is not always optimal as both position and item dependencies result in a complex reward function. A very naive example is the lack of diversity when all the most attractive items are from the same category. We model the position and item dependencies in the ordered list and design UCB and Thompson Sampling type algorithms for this problem. Our work generalizes existing studies in several directions, including position dependencies where position discount is a particular case, and connecting the ranking problem to graph theory.
翻译:我们研究广义线性Bandits中的排序问题。在每个时间步,学习代理选择有序项目列表并观测随机结果。在推荐系统中,显示最具吸引力项目的有序列表并非始终最优,因为位置依赖和项目依赖导致复杂的奖励函数。一个非常简单的例子是当所有最具吸引力的项目都来自同一类别时会出现多样性缺失。我们为有序列表中的位置依赖和项目依赖建模,并针对该问题设计了置信界上界(UCB)和汤普森采样(Thompson Sampling)类型算法。本研究在多个方向对现有成果进行了推广,包括将位置折扣作为特例的位置依赖建模,以及将排序问题与图论建立关联。