Recommendation systems are a core feature of social media companies with their uses including recommending organic and promoted contents. Many modern recommendation systems are split into multiple stages - candidate generation and heavy ranking - to balance computational cost against recommendation quality. We focus on the candidate generation phase of a large-scale ads recommendation problem in this paper, and present a machine learning first heterogeneous re-architecture of this stage which we term TwERC. We show that a system that combines a real-time light ranker with sourcing strategies capable of capturing additional information provides validated gains. We present two strategies. The first strategy uses a notion of similarity in the interaction graph, while the second strategy caches previous scores from the ranking stage. The graph based strategy achieves a 4.08% revenue gain and the rankscore based strategy achieves a 1.38% gain. These two strategies have biases that complement both the light ranker and one another. Finally, we describe a set of metrics that we believe are valuable as a means of understanding the complex product trade offs inherent in industrial candidate generation systems.
翻译:推荐系统是社交媒体公司的核心功能,其应用包括推荐自然内容和推广内容。许多现代推荐系统分为多个阶段——候选生成和深度排序,以在计算成本与推荐质量之间取得平衡。本文聚焦于大规模广告推荐问题中的候选生成阶段,并提出一种以机器学习为先的异构重构方案,我们称之为TwERC。我们展示了一种结合实时轻量排序器与能够捕获额外信息的源策略的系统,可提供已验证的性能提升。我们提出两种策略:第一种策略利用交互图中的相似性概念,第二种策略缓存来自排序阶段的先前得分。基于图的策略实现了4.08%的收入增长,而基于排序得分的策略实现了1.38%的增长。这两种策略的偏差相互补充,且与轻量排序器形成互补。最后,我们描述了一组指标,我们认为这些指标对于理解工业级候选生成系统中固有的复杂产品权衡具有重要价值。