Effective item identifiers (IDs) are an important component for recommender systems (RecSys) in practice, and are commonly adopted in many use cases such as retrieval and ranking. IDs can encode collaborative filtering signals within training data, such that RecSys models can extrapolate during the inference and personalize the prediction based on users' behavioral histories. Recently, Semantic IDs (SIDs) have become a trending paradigm for RecSys. In comparison to the conventional atomic ID, an SID is an ordered list of codes, derived from tokenizers such as residual quantization, applied to semantic representations commonly extracted from foundation models or collaborative signals. SIDs have drastically smaller cardinality than the atomic counterpart, and induce semantic clustering in the ID space. At Snapchat, we apply SIDs as auxiliary features for ranking models, and also explore SIDs as additional retrieval sources in different ML applications. In this paper, we discuss practical technical challenges we encountered while applying SIDs, experiments we have conducted, and design choices we have iterated to mitigate these challenges. Backed by promising offline results on both internal data and academic benchmarks as well as online A/B studies, SID variants have been launched in multiple production models with positive metrics impact.
翻译:有效的物品标识符(ID)是推荐系统(RecSys)在实际应用中的重要组成部分,广泛应用于检索和排序等多种场景。ID能够编码训练数据中的协同过滤信号,使推荐系统在推理过程中进行外推,并根据用户行为历史实现个性化预测。近年来,语义ID(SID)已成为推荐系统的新兴范式。与传统原子ID相比,SID是一种有序编码序列,由残差量化等分词器从通常从基础模型或协同信号中提取的语义表示中衍生而来。SID的基数远小于原子ID,并在ID空间中诱导语义聚类。在Snapchat,我们将SID作为排序模型的辅助特征,并探索将SID作为不同机器学习应用中的额外检索来源。本文讨论了我们在应用SID时遇到的实际技术挑战、进行的实验以及为缓解这些挑战而迭代的设计选择。基于内部数据和学术基准上令人鼓舞的离线结果以及在线A/B实验,SID变体已在多个生产模型中上线,并取得了积极的指标影响。