We study the problem of optimizing a recommender system for outcomes that occur over several weeks or months. We begin by drawing on reinforcement learning to formulate a comprehensive model of users' recurring relationships with a recommender system. Measurement, attribution, and coordination challenges complicate algorithm design. We describe careful modeling -- including a new representation of user state and key conditional independence assumptions -- which overcomes these challenges and leads to simple, testable recommender system prototypes. We apply our approach to a podcast recommender system that makes personalized recommendations to hundreds of millions of listeners. A/B tests demonstrate that purposefully optimizing for long-term outcomes leads to large performance gains over conventional approaches that optimize for short-term proxies.
翻译:我们研究了优化推荐系统以在数周或数月内实现长期效果的问题。首先,我们借鉴强化学习方法,构建了一个描述用户与推荐系统之间周期性关系的全面模型。测量、归因与协同调度方面的挑战使算法设计复杂化。我们提出了一套精细的建模方案——包括新的用户状态表示及关键的条件独立性假设——这些方案克服了上述挑战,并催生了简单、可测试的推荐系统原型。我们将该方法应用于面向数亿听众进行个性化推荐的播客推荐系统。A/B测试表明,与传统以短期代理指标为优化目标的方案相比,有意针对长期效果进行优化可带来显著的性能提升。