Multi-class classification methods that produce sets of probabilistic classifiers, such as ensemble learning methods, are able to model aleatoric and epistemic uncertainty. Aleatoric uncertainty is then typically quantified via the Bayes error, and epistemic uncertainty via the size of the set. In this paper, we extend the notion of calibration, which is commonly used to evaluate the validity of the aleatoric uncertainty representation of a single probabilistic classifier, to assess the validity of an epistemic uncertainty representation obtained by sets of probabilistic classifiers. Broadly speaking, we call a set of probabilistic classifiers calibrated if one can find a calibrated convex combination of these classifiers. To evaluate this notion of calibration, we propose a novel nonparametric calibration test that generalizes an existing test for single probabilistic classifiers to the case of sets of probabilistic classifiers. Making use of this test, we empirically show that ensembles of deep neural networks are often not well calibrated.
翻译:多类分类方法(如集成学习方法)可生成概率分类器集合,从而对偶然不确定性和认知不确定性进行建模。通常,偶然不确定性通过贝叶斯误差量化,而认知不确定性则通过集合的规模度量。本文扩展了校准的概念——该概念常用于评估单一概率分类器偶然不确定性表示的有效性——以评估概率分类器集合所获得的认知不确定性表示的有效性。广义而言,若能从一组概率分类器中找到校准的凸组合,则称该集合为校准的。为评估这一校准概念,我们提出了一种新型非参数校准检验方法,将现有针对单一概率分类器的检验推广至概率分类器集合的情况。利用该检验,我们通过实验表明,深度神经网络的集成通常未得到良好校准。