The practice of uncertainty quantification (UQ) validation, notably in machine learning for the physico-chemical sciences, rests on several graphical methods (scattering plots, calibration curves, reliability diagrams and confidence curves) which explore complementary aspects of calibration, without covering all the desirable ones. For instance, none of these methods deals with the reliability of UQ metrics across the range of input features. Based on three complementary concepts, calibration, consistency and adaptivity, the toolbox of common validation methods for variance- and intervals- based metrics is revisited with the aim to provide a better grasp on their capabilities. This study is conceived as an introduction to UQ validation, and all methods are derived from a few basic rules. The methods are illustrated and tested on synthetic datasets and examples extracted from the recent physico-chemical machine learning UQ literature.
翻译:不确定性量化(UQ)验证实践,特别是在物理化学科学领域的机器学习中,依赖于几种图形方法(散点图、校准曲线、可靠性图和置信曲线),这些方法探索了校准的互补方面,但未能涵盖所有理想特性。例如,这些方法均未涉及UQ度量在输入特征范围内的可靠性。基于校准、一致性和自适应性这三个互补概念,本文重新审视了基于方差和区间的常见验证方法工具箱,旨在更清晰地把握其能力。本研究作为UQ验证的入门介绍,所有方法均源自少数基本规则。这些方法通过合成数据集以及从近期物理化学机器学习UQ文献中提取的示例进行了说明和测试。