Approaches to recommendation are typically evaluated in one of two ways: (1) via a (simulated) online experiment, often seen as the gold standard, or (2) via some offline evaluation procedure, where the goal is to approximate the outcome of an online experiment. Several offline evaluation metrics have been adopted in the literature, inspired by ranking metrics prevalent in the field of Information Retrieval. (Normalised) Discounted Cumulative Gain (nDCG) is one such metric that has seen widespread adoption in empirical studies, and higher (n)DCG values have been used to present new methods as the state-of-the-art in top-$n$ recommendation for many years. Our work takes a critical look at this approach, and investigates when we can expect such metrics to approximate the gold standard outcome of an online experiment. We formally present the assumptions that are necessary to consider DCG an unbiased estimator of online reward and provide a derivation for this metric from first principles, highlighting where we deviate from its traditional uses in IR. Importantly, we show that normalising the metric renders it inconsistent, in that even when DCG is unbiased, ranking competing methods by their normalised DCG can invert their relative order. Through a correlation analysis between off- and on-line experiments conducted on a large-scale recommendation platform, we show that our unbiased DCG estimates strongly correlate with online reward, even when some of the metric's inherent assumptions are violated. This statement no longer holds for its normalised variant, suggesting that nDCG's practical utility may be limited.
翻译:推荐方法通常通过以下两种方式之一进行评估:(1) 通过(模拟)在线实验(通常被视为黄金标准),或 (2) 通过某种离线评估程序,其目标是近似在线实验的结果。受信息检索领域流行的排序指标启发,文献中已采用了几种离线评估指标。(归一化)折损累计增益 (nDCG) 就是这样一种在实证研究中被广泛采用的指标,多年来,更高的 (n)DCG 值一直被用于将新方法展示为 Top-$n$ 推荐中的最先进技术。我们的工作对此方法进行了批判性审视,并研究了在何种情况下可以期望此类指标近似于在线实验的黄金标准结果。我们正式提出了将 DCG 视为在线奖励无偏估计量所必需的假设,并从基本原理出发推导了该指标,突出了我们偏离其在信息检索中传统用途的地方。重要的是,我们表明对指标进行归一化会使其不一致,即即使 DCG 是无偏的,通过归一化 DCG 对竞争方法进行排序也可能颠倒它们的相对顺序。通过在一个大规模推荐平台上进行的离线和在线实验之间的相关性分析,我们表明即使违反了该指标的某些固有假设,我们的无偏 DCG 估计也与在线奖励强相关。这一结论不再适用于其归一化变体,这表明 nDCG 的实际效用可能有限。