Recent research has suggested different metrics to measure the inconsistency of recommendation performance, including the accuracy difference between user groups, miscalibration, and popularity lift. However, a study that relates miscalibration and popularity lift to recommendation accuracy across different user groups is still missing. Additionally, it is unclear if particular genres contribute to the emergence of inconsistency in recommendation performance across user groups. In this paper, we present an analysis of these three aspects of five well-known recommendation algorithms for user groups that differ in their preference for popular content. Additionally, we study how different genres affect the inconsistency of recommendation performance, and how this is aligned with the popularity of the genres. Using data from LastFm, MovieLens, and MyAnimeList, we present two key findings. First, we find that users with little interest in popular content receive the worst recommendation accuracy, and that this is aligned with miscalibration and popularity lift. Second, our experiments show that particular genres contribute to a different extent to the inconsistency of recommendation performance, especially in terms of miscalibration in the case of the MyAnimeList dataset.
翻译:近期研究提出了多种衡量推荐性能不一致性的指标,包括用户群体间的准确率差异、校准偏差以及流行度提升。然而,目前尚缺乏将校准偏差和流行度提升与不同用户群体的推荐准确性相关联的研究。此外,特定内容类型是否会导致不同用户群体间推荐性能的不一致性仍不明确。本文针对偏好流行内容的程度不同的用户群体,分析了五种主流推荐算法在上述三个维度上的表现。同时,我们研究了不同内容类型如何影响推荐性能的不一致性,以及这种不一致性与内容类型流行度的关系。基于LastFm、MovieLens和MyAnimeList数据集,我们得出两个关键发现:第一,对流行内容兴趣较低的用户群体获得的推荐准确率最差,且这一现象与校准偏差和流行度提升具有一致性;第二,实验表明特定内容类型对不同用户群体间推荐性能不一致性的贡献程度存在差异,尤其在MyAnimeList数据集中表现为校准偏差的显著差异。