We generalize fast Gaussian process leave-one-out formulae to multiple-fold cross-validation, highlighting in turn the covariance structure of cross-validation residuals in both Simple and Universal Kriging frameworks. We illustrate how resulting covariances affect model diagnostics. We further establish in the case of noiseless observations that correcting for covariances between residuals in cross-validation-based estimation of the scale parameter leads back to MLE. Also, we highlight in broader settings how differences between pseudo-likelihood and likelihood methods boil down to accounting or not for residual covariances. The proposed fast calculation of cross-validation residuals is implemented and benchmarked against a naive implementation. Numerical experiments highlight the accuracy and substantial speed-ups that our approach enables. However, as supported by a discussion on main drivers of computational costs and by a numerical benchmark, speed-ups steeply decline as the number of folds (say, all sharing the same size) decreases. An application to a contaminant localization test case illustrates that grouping clustered observations in folds may help improving model assessment and parameter fitting compared to Leave-One-Out. Overall, our results enable fast multiple-fold cross-validation, have direct consequences in model diagnostics, and pave the way to future work on hyperparameter fitting and on the promising field of goal-oriented fold design.
翻译:我们将快速高斯过程留一法公式推广到多折交叉验证,重点阐明了简单克里金和通用克里金框架中交叉验证残差的协方差结构。我们展示了所得协方差如何影响模型诊断。进一步,在无噪声观测情形中,我们证实了基于交叉验证的尺度参数估计中校正残差间协方差会重新得到最大似然估计。此外,我们在更广泛背景下阐明了伪似然方法与似然方法之间的差异可归结于是否考虑残差协方差。所提出的交叉验证残差快速计算方法已实现,并与朴素实现进行了基准测试。数值实验凸显了我们方法所实现的精度和显著加速效果。然而,对计算成本主要驱动因素的讨论及数值基准测试表明,随着折数(假设所有折规模相同)减少,加速效果急剧下降。应用于污染物定位测试案例表明,与留一法相比,将聚类观测分组到各折中有助于改进模型评估和参数拟合。总体而言,我们的结果实现了快速多折交叉验证,对模型诊断具有直接影响,并为超参数拟合及目标导向折设计这一前景广阔领域的未来研究铺平了道路。