Leave-one-out cross-validation (LOO-CV) is a popular method for comparing Bayesian models based on their estimated predictive performance on new, unseen, data. As leave-one-out cross-validation is based on finite observed data, there is uncertainty about the expected predictive performance on new data. By modeling this uncertainty when comparing two models, we can compute the probability that one model has a better predictive performance than the other. Modeling this uncertainty well is not trivial, and for example, it is known that the commonly used standard error estimate is often too small. We study the properties of the Bayesian LOO-CV estimator and the related uncertainty estimates when comparing two models. We provide new results of the properties both theoretically in the linear regression case and empirically for multiple different models and discuss the challenges of modeling the uncertainty. We show that problematic cases include: comparing models with similar predictions, misspecified models, and small data. In these cases, there is a weak connection in the skewness of the individual leave-one-out terms and the distribution of the error of the Bayesian LOO-CV estimator. We show that it is possible that the problematic skewness of the error distribution, which occurs when the models make similar predictions, does not fade away when the data size grows to infinity in certain situations. Based on the results, we also provide practical recommendations for the users of Bayesian LOO-CV for model comparison.
翻译:留一交叉验证(LOO-CV)是一种基于模型在新观测数据上的预测性能估计来比较贝叶斯模型的常用方法。由于留一交叉验证基于有限的观测数据,因此对新数据期望预测性能的估计存在不确定性。通过建模这种不确定性,我们可以在比较两个模型时计算其中一个模型预测性能优于另一个的概率。准确建模这种不确定性并非易事,例如,已知常用的标准误差估计往往偏小。本文研究了贝叶斯LOO-CV估计量及其相关不确定性估计在比较两个模型时的性质。我们在线性回归情形下从理论上给出了性质的新结果,并针对多种不同模型进行了实证分析,同时讨论了建模不确定性的挑战。研究表明,问题情形包括:比较预测相似的模型、模型误设以及小样本数据。在这些情形下,个体留一项的偏度与贝叶斯LOO-CV估计量误差分布之间存在较弱的联系。我们证明,当模型预测相似时,误差分布的问题性偏度在某些情况下不会随着数据规模无限增长而消失。基于这些结果,我们还为使用贝叶斯LOO-CV进行模型比较的用户提供了实用建议。