We show that the effectiveness of the well celebrated Mixup [Zhang et al., 2018] can be further improved if instead of using it as the sole learning objective, it is utilized as an additional regularizer to the standard cross-entropy loss. This simple change not only provides much improved accuracy but also significantly improves the quality of the predictive uncertainty estimation of Mixup in most cases under various forms of covariate shifts and out-of-distribution detection experiments. In fact, we observe that Mixup yields much degraded performance on detecting out-of-distribution samples possibly, as we show empirically, because of its tendency to learn models that exhibit high-entropy throughout; making it difficult to differentiate in-distribution samples from out-distribution ones. To show the efficacy of our approach (RegMixup), we provide thorough analyses and experiments on vision datasets (ImageNet & CIFAR-10/100) and compare it with a suite of recent approaches for reliable uncertainty estimation.
翻译:我们证明,若将广受赞誉的Mixup方法[Zhang等人,2018]作为标准交叉熵损失的附加正则化器而非唯一学习目标,其有效性可进一步提升。这一简单改进不仅显著提高了准确率,更在多数协变量偏移与分布外检测实验场景中,大幅改善了Mixup预测不确定性估计的质量。事实上,我们观察到Mixup在检测分布外样本时性能显著下降——如经验分析所示,这可能源于其倾向于学习具有全局高熵特征的模型,导致难以区分分布内样本与分布外样本。为验证本方法(RegMixup)的有效性,我们在视觉数据集(ImageNet与CIFAR-10/100)上进行了详尽分析与实验,并与近期一系列面向可靠不确定性估计的方法进行了对比。