Rank and PIT histograms are established tools to assess the calibration of probabilistic forecasts. They not only check whether an ensemble forecast is calibrated, but they also reveal what systematic biases (if any) are present in the forecasts. Several extensions of rank histograms have been proposed to evaluate the calibration of probabilistic forecasts for multivariate outcomes. These extensions introduce a so-called pre-rank function that condenses the multivariate forecasts and observations into univariate objects, from which a standard rank histogram can be produced. Existing pre-rank functions typically aim to preserve as much information as possible when condensing the multivariate forecasts and observations into univariate objects. Although this is sensible when conducting statistical tests for multivariate calibration, it can hinder the interpretation of the resulting histograms. In this paper, we demonstrate that there are few restrictions on the choice of pre-rank function, meaning forecasters can choose a pre-rank function depending on what information they want to extract from their forecasts. We introduce the concept of simple pre-rank functions, and provide examples that can be used to assess the location, scale, and dependence structure of multivariate probabilistic forecasts, as well as pre-rank functions tailored to the evaluation of probabilistic spatial field forecasts. The simple pre-rank functions that we introduce are easy to interpret, easy to implement, and they deliberately provide complementary information, meaning several pre-rank functions can be employed to achieve a more complete understanding of multivariate forecast performance. We then discuss how e-values can be employed to formally test for multivariate calibration over time. This is demonstrated in an application to wind speed forecasting using the EUPPBench post-processing benchmark data set.
翻译:秩和PIT直方图是评估概率预报校准能力的成熟工具。它们不仅能检验集合预报是否校准,还能揭示预报中存在的系统性偏差(如有)。针对多变量结果的概率预报校准评估,研究者提出了多种秩直方图扩展方法。这些扩展引入了所谓的预秩函数,将多变量预报和观测值压缩为单变量对象,从而生成标准秩直方图。现有预秩函数通常旨在压缩多变量预报和观测值时尽可能保留信息。尽管这对多变量校准的统计检验是合理的,但可能阻碍对生成直方图的解释。本文表明,预秩函数的选择几乎没有限制,这意味着预报员可根据需要提取的信息选择预秩函数。我们引入简单预秩函数的概念,并提供可评估多变量概率预报位置、尺度及依赖结构的示例,以及针对概率空间场预报评估定制的预秩函数。我们提出的简单预秩函数易于解释、易于实现,且有意提供互补信息,即可通过多种预秩函数更全面地理解多变量预报性能。随后讨论如何利用e值对时间序列上的多变量校准进行形式化检验。这一方法通过EUPPBench后处理基准数据集在风速预报中的应用得到验证。