In an era of countless content offerings, recommender systems alleviate information overload by providing users with personalized content suggestions. Due to the scarcity of explicit user feedback, modern recommender systems typically optimize for the same fixed combination of implicit feedback signals across all users. However, this approach disregards a growing body of work highlighting that (i) implicit signals can be used by users in diverse ways, signaling anything from satisfaction to active dislike, and (ii) different users communicate preferences in different ways. We propose applying the recent Interaction Grounded Learning (IGL) paradigm to address the challenge of learning representations of diverse user communication modalities. Rather than requiring a fixed, human-designed reward function, IGL is able to learn personalized reward functions for different users and then optimize directly for the latent user satisfaction. We demonstrate the success of IGL with experiments using simulations as well as with real-world production traces.
翻译:在内容供给过载的时代,推荐系统通过向用户提供个性化内容建议来缓解信息过载问题。由于显式用户反馈的稀缺性,现代推荐系统通常对所有用户优化相同的固定隐式反馈信号组合。然而,这种方法忽视了越来越多研究强调的两点事实:(i)隐式信号可能被用户以多样化方式使用,其含义从满意到明显厌恶不等;(ii)不同用户表达偏好的方式存在差异。本文提出应用近期发展的交互式学习(IGL)范式,以解决不同用户通信模式表征学习的挑战。IGL无需固定的、人工设计的奖励函数,而是能够为不同用户学习个性化奖励函数,并直接针对潜在用户满意度进行优化。我们通过仿真实验和真实生产环境数据验证了IGL的成功。