Large language models (LLMs) have gained much attention in the recommendation community; some studies have observed that LLMs, fine-tuned by the cross-entropy loss with a full softmax, could achieve state-of-the-art performance already. However, these claims are drawn from unobjective and unfair comparisons. In view of the substantial quantity of items in reality, conventional recommenders typically adopt a pointwise/pairwise loss function instead for training. This substitute however causes severe performance degradation, leading to under-estimation of conventional methods and over-confidence in the ranking capability of LLMs. In this work, we theoretically justify the superiority of cross-entropy, and showcase that it can be adequately replaced by some elementary approximations with certain necessary modifications. The remarkable results across three public datasets corroborate that even in a practical sense, existing LLM-based methods are not as effective as claimed for next-item recommendation. We hope that these theoretical understandings in conjunction with the empirical results will facilitate an objective evaluation of LLM-based recommendation in the future.
翻译:大语言模型(LLMs)在推荐领域引起了广泛关注;一些研究发现,通过完整softmax的交叉熵损失进行微调的大语言模型已能达到最优性能。然而,这些结论源于不客观且不公平的对比。鉴于现实中商品数量巨大,传统推荐系统通常采用逐点/逐对损失函数进行训练。但这种替代方式会导致性能严重下降,既低估了传统方法,又过度信任大语言模型的排序能力。本研究从理论上论证了交叉熵的优越性,并证明通过必要的改进,某些基础近似方法可以充分替代交叉熵损失。在三组公开数据集上的显著结果证实,即使在实际应用中,现有基于大语言模型的方法在下一项推荐任务中并未达到宣称的效果。我们期望这些理论认知与实证结果的结合,能为未来大语言模型推荐系统提供客观评估依据。