A key puzzle in search, ads, and recommendation is that the ranking model can only utilize a small portion of the vastly available user interaction data. As a result, increasing data volume, model size, or computation FLOPs will quickly suffer from diminishing returns. We examined this problem and found that one of the root causes may lie in the so-called ``item-centric'' formulation, which has an unbounded vocabulary and thus uncontrolled model complexity. To mitigate quality saturation, we introduce an alternative formulation named ``user-centric ranking'', which is based on a transposed view of the dyadic user-item interaction data. We show that this formulation has a promising scaling property, enabling us to train better-converged models on substantially larger data sets.
翻译:搜索、广告和推荐面临的一个核心难题是:排序模型只能利用海量用户交互数据中的一小部分。因此,增加数据量、模型规模或计算FLOPs会迅速遭遇收益递减。我们研究了这一问题,发现其根本原因之一可能在于所谓的“以物品为中心”的公式化方法,该方法具有无界词表,因而模型复杂度不可控。为了缓解质量饱和,我们提出了一种替代性公式化方法,名为“以用户为中心的排序”,该方法基于对用户-物品二元交互数据的转置视角。我们证明,该公式具有有前景的扩展特性,使我们能够在规模大得多的数据集上训练出收敛更好的模型。