Computer vision systems that are deployed in safety-critical applications need to quantify their output uncertainty. We study regression from images to parameter values and here it is common to detect uncertainty by predicting probability distributions. In this context, we investigate the regression-by-classification paradigm which can represent multimodal distributions, without a prior assumption on the number of modes. Through experiments on a specifically designed synthetic dataset, we demonstrate that traditional loss functions lead to poor probability distribution estimates and severe overconfidence, in the absence of full ground truth distributions. In order to alleviate these issues, we propose hinge-Wasserstein -- a simple improvement of the Wasserstein loss that reduces the penalty for weak secondary modes during training. This enables prediction of complex distributions with multiple modes, and allows training on datasets where full ground truth distributions are not available. In extensive experiments, we show that the proposed loss leads to substantially better uncertainty estimation on two challenging computer vision tasks: horizon line detection and stereo disparity estimation.
翻译:部署在安全关键应用中的计算机视觉系统需要量化其输出不确定性。我们研究从图像到参数值的回归问题,在此类问题中,通常通过预测概率分布来检测不确定性。在此背景下,我们探究了无需预先假设模态数量的回归-分类范式,该范式能够表示多模态分布。通过在专门设计的合成数据集上的实验,我们证明:在缺乏完整真实分布的情况下,传统损失函数会导致概率分布估计质量低下和严重的过度自信。为缓解这些问题,我们提出铰链-沃瑟斯坦(Hinge-Wasserstein)——这是对沃瑟斯坦损失的简单改进,可在训练过程中降低对次要弱模态的惩罚。这使得模型能够预测包含多个模态的复杂分布,并允许在缺乏完整真实分布的数据集上进行训练。在大量实验中,我们表明所提出的损失函数在两项具有挑战性的计算机视觉任务(地平线检测和立体视差估计)上显著提升了不确定性估计质量。