Some popular Machine Learning Uncertainty Quantification (ML-UQ) calibration statistics do not have predefined reference values and are mostly used in comparative studies. In consequence, calibration is almost never validated and the diagnostic is left to the appreciation of the reader. Simulated reference values, based on synthetic calibrated datasets derived from actual uncertainties, have been proposed to palliate this problem. As the generative probability distribution for the simulation of synthetic errors is often not constrained, the sensitivity of simulated reference values to the choice of generative distribution might be problematic, shedding a doubt on the calibration diagnostic. This study explores various facets of this problem, and shows that some statistics are excessively sensitive to the choice of generative distribution to be used for validation when the generative distribution is unknown. This is the case, for instance, of the correlation coefficient between absolute errors and uncertainties (CC) and of the expected normalized calibration error (ENCE). A robust validation workflow to deal with simulated reference values is proposed.
翻译:部分流行的机器学习不确定性量化(ML-UQ)校准统计量缺乏预定义的参考值,主要被用于比较性研究。因此,校准几乎从未得到验证,诊断结果只能交由读者自行判断。为解决这一问题,研究者提出了基于实际不确定性生成的合成校准数据集模拟参考值。由于合成误差的生成概率分布通常不受约束,模拟参考值对生成分布选择的敏感性可能带来问题,从而对校准诊断产生疑虑。本研究探讨了该问题的多个方面,并表明当生成分布未知时,部分统计量对用于验证的生成分布选择过于敏感——例如,绝对误差与不确定性的相关系数(CC)以及期望归一化校准误差(ENCE)即属于此类情况。本文提出了一种适用于模拟参考值的稳健验证工作流程。