The evaluation of recommender systems from a practical perspective is a topic of ongoing discourse within the research community. While many current evaluation methods reduce performance to a single value metric as an easy way to compare models, it relies on the assumption that the methods' performance remains constant over time. In this study, we examine this assumption and propose the Cross-Validation Thought Time (CVTT) technique as a more comprehensive evaluation method, focusing on model performance over time. By utilizing the proposed technique, we conduct an in-depth analysis of the performance of popular RecSys algorithms. Our findings indicate that (1) the performance of the recommenders varies over time for all reviewed datasets, (2) using simple evaluation approaches can lead to a substantial decrease in performance in real-world evaluation scenarios, and (3) excessive data usage can lead to suboptimal results.
翻译:从实践角度评估推荐系统是研究社区中持续讨论的话题。尽管当前许多评估方法将性能简化为单一指标值,以便于模型比较,但这依赖于模型性能随时间保持恒定的假设。本研究考察了这一假设,并提出时间交叉验证(CVTT)技术作为更全面的评估方法,重点关注模型随时间变化的性能。通过运用该技术,我们对主流推荐系统算法的性能进行了深入分析。研究结果表明:(1)在所有考察的数据集中,推荐器的性能随时间呈现波动;(2)使用简单评估方法会导致实际场景中的性能显著下降;(3)过度使用数据可能导致次优结果。