We propose a novel, succinct, and effective approach for distribution prediction to quantify uncertainty in machine learning. It incorporates adaptively flexible distribution prediction of $\mathbb{P}(\mathbf{y}|\mathbf{X}=x)$ in regression tasks. This conditional distribution's quantiles of probability levels spreading the interval $(0,1)$ are boosted by additive models which are designed by us with intuitions and interpretability. We seek an adaptive balance between the structural integrity and the flexibility for $\mathbb{P}(\mathbf{y}|\mathbf{X}=x)$, while Gaussian assumption results in a lack of flexibility for real data and highly flexible approaches (e.g., estimating the quantiles separately without a distribution structure) inevitably have drawbacks and may not lead to good generalization. This ensemble multi-quantiles approach called EMQ proposed by us is totally data-driven, and can gradually depart from Gaussian and discover the optimal conditional distribution in the boosting. On extensive regression tasks from UCI datasets, we show that EMQ achieves state-of-the-art performance comparing to many recent uncertainty quantification methods. Visualization results further illustrate the necessity and the merits of such an ensemble model.
翻译:我们提出一种新颖、简洁且有效的分布预测方法,用于量化机器学习中的不确定性。该方法在回归任务中融入对 $\mathbb{P}(\mathbf{y}|\mathbf{X}=x)$ 的自适应灵活分布预测。通过我们设计、具有直观性和可解释性的加性模型,增强覆盖区间 $(0,1)$ 的概率水平下的条件分布分位数。我们在 $\mathbb{P}(\mathbf{y}|\mathbf{X}=x)$ 的结构完整性与灵活性之间寻求自适应平衡,而高斯假设会导致真实数据缺乏灵活性,高度灵活的方法(例如,无分布结构地单独估计分位数)则不可避免地存在缺陷且可能无法实现良好的泛化。我们提出的这种称为EMQ的集成多分位数方法完全由数据驱动,可在提升过程中逐步偏离高斯分布并发现最优条件分布。针对UCI数据集上的大量回归任务,我们表明,与众多近期不确定性量化方法相比,EMQ实现了最先进的性能。可视化结果进一步阐明了此类集成模型的必要性与优势。