A widely held hypothesis for why generative recommendation (GR) models outperform conventional item ID-based models is that they generalize better. However, there is few systematic way to verify this hypothesis beyond a superficial comparison of overall performance. To address this gap, we categorize each data instance based on the specific capability required for a correct prediction: either memorization (reusing item transition patterns observed during training) or generalization (composing known patterns to predict unseen item transitions). Extensive experiments show that GR models perform better on instances that require generalization, whereas item ID-based models perform better when memorization is more important. To explain this divergence, we shift the analysis from the item level to the token level and show that what appears to be item-level generalization often reduces to token-level memorization for GR models. Finally, we show that the two paradigms are complementary. We propose a simple memorization-aware indicator that adaptively combines them on a per-instance basis, leading to improved overall recommendation performance.
翻译:广泛认为,生成式推荐(GR)模型优于传统基于项目ID的模型的一个核心原因是其泛化能力更强。然而,现有研究大多仅通过整体性能的表面比较来验证这一假设,缺乏系统性的分析。为弥补这一空白,我们根据正确预测所需的具体能力(即:记忆能力——复用训练中观察到的项目转移模式,或泛化能力——组合已知模式推测未见过的项目转移)对每个数据实例进行分类。大量实验表明,GR模型在处理需要泛化能力的实例时表现更优,而基于项目ID的模型则在更依赖记忆能力的场景中更具优势。为解释这一差异,我们将分析视角从项目层级降至词元层级,发现GR模型中看似项目级别的泛化能力,实则常可归约为词元级别的记忆能力。此外,我们揭示两种范式具有互补性,并提出一种简单的记忆感知指标,通过逐实例自适应的方式融合二者,从而显著提升整体推荐性能。