Large-scale pre-trained models have achieved remarkable success in a variety of scenarios and applications, but how to leverage them to improve the prediction reliability of downstream models is undesirably under-explored. Moreover, modern neural networks have been found to be poorly calibrated and make overconfident predictions regardless of inherent sample difficulty and data uncertainty. To address this issue, we propose to utilize large-scale pre-trained models to guide downstream model training with sample difficulty-aware entropy regularization. Pre-trained models that have been exposed to large-scale datasets and do not overfit the downstream training classes enable us to measure each training sample difficulty via feature-space Gaussian modeling and relative Mahalanobis distance computation. Importantly, by adaptively penalizing overconfident prediction based on the sample's difficulty, we simultaneously improve accuracy and uncertainty calibration on various challenging benchmarks, consistently surpassing competitive baselines for reliable prediction.
翻译:大规模预训练模型已在多种场景和应用中取得了显著成功,然而如何利用它们提升下游模型的预测可靠性却未得到充分探索。此外,现代神经网络被发现存在校准不良的问题,且无论样本固有难度与数据不确定性如何,都会做出过度自信的预测。针对这一问题,我们提出利用大规模预训练模型,通过引入基于样本难度感知的熵正则化来指导下游模型训练。在大规模数据集上训练且未对下游训练类别产生过拟合的预训练模型,使我们能够通过特征空间的高斯建模和相对马氏距离计算,衡量每个训练样本的难度。重要的是,通过根据样本难度自适应地惩罚过度自信的预测,我们在多个具有挑战性的基准测试上同时提升了准确率与不确定性校准效果,在可靠预测方面持续超越竞争基线方法。