As neural networks become more popular, the need for accompanying uncertainty estimates increases. There are currently two main approaches to test the quality of these estimates. Most methods output a density. They can be compared by evaluating their loglikelihood on a test set. Other methods output a prediction interval directly. These methods are often tested by examining the fraction of test points that fall inside the corresponding prediction intervals. Intuitively both approaches seem logical. However, we demonstrate through both theoretical arguments and simulations that both ways of evaluating the quality of uncertainty estimates have serious flaws. Firstly, both approaches cannot disentangle the separate components that jointly create the predictive uncertainty, making it difficult to evaluate the quality of the estimates of these components. Secondly, a better loglikelihood does not guarantee better prediction intervals, which is what the methods are often used for in practice. Moreover, the current approach to test prediction intervals directly has additional flaws. We show why it is fundamentally flawed to test a prediction or confidence interval on a single test set. At best, marginal coverage is measured, implicitly averaging out overconfident and underconfident predictions. A much more desirable property is pointwise coverage, requiring the correct coverage for each prediction. We demonstrate through practical examples that these effects can result in favoring a method, based on the predictive uncertainty, that has undesirable behaviour of the confidence or prediction intervals. Finally, we propose a simulation-based testing approach that addresses these problems while still allowing easy comparison between different methods.
翻译:随着神经网络日益普及,对其伴随的不确定性估计的需求也随之增加。目前主要有两种方法来测试这些估计的质量。大多数方法输出一个密度分布,可以通过评估其在测试集上的对数似然来进行比较。另一些方法则直接输出预测区间,通常通过检查落于对应预测区间内的测试点比例来测试。直觉上,这两种方法似乎都合乎逻辑。然而,我们通过理论论证和模拟实验证明,这两种评估不确定性估计质量的方法都存在严重缺陷。首先,这两种方法都无法分离共同构成预测不确定性的各个独立分量,从而难以评估这些分量估计的质量。其次,更好的对数似然并不能保证更好的预测区间,而预测区间恰恰是这些方法在实际中常用的目的。此外,当前直接测试预测区间的方法还存在其他缺陷。我们论证了在单一测试集上测试预测区间或置信区间在根本上就是有缺陷的。这种方法最多只能衡量边际覆盖,它隐式地平均了过度自信和信心不足的预测。一种更理想的属性是逐点覆盖,即要求每个预测都具有正确的覆盖概率。我们通过实际例子证明,这些效应可能导致基于预测不确定性而偏向某个方法,而该方法本身却具有不良的置信区间或预测区间行为。最后,我们提出了一种基于模拟的测试方法,该方法解决了上述问题,同时仍允许在不同方法之间进行简便比较。