Large pre-trained language models (PLMs) have demonstrated strong performance on natural language understanding (NLU) tasks through fine-tuning. However, fine-tuned models still suffer from overconfident predictions, especially in out-of-domain settings. In this paper, we tackle the problem of calibrating fine-tuned language models. We demonstrate that the PLMs are well-calibrated on the masked language modeling task with robust predictive confidence under domain shift, yet the fine-tuned models fail to retain such property due to catastrophic forgetting, which impacts the calibration on the downstream classification task. In light of these observations, we evaluate the calibration of several methods that preserve pre-trained features and show that preserving pre-trained features can improve the calibration of fine-tuned language models. Among these methods, our proposed method that encourages the fine-tuned model to learn generative representations with auxiliary language modeling objective achieves competitive accuracy and the lowest expected calibration error compared to several strong baselines under both in-domain and out-of-domain settings on three downstream NLU tasks.
翻译:大型预训练语言模型(PLMs)通过微调在自然语言理解(NLU)任务中展现出强大性能。然而,微调后的模型仍存在过度自信预测的问题,尤其在域外场景中。本文致力于解决微调语言模型的校准问题。我们证明,PLMs在掩码语言建模任务上具有良好的校准性,且在域偏移下仍保持稳健的预测置信度,但微调模型因灾难性遗忘而无法保留这一特性,从而影响下游分类任务的校准。基于这些发现,我们评估了多种保留预训练特征的方法的校准性能,结果表明:保留预训练特征能够改善微调语言模型的校准效果。在这些方法中,我们提出的方法通过辅助语言建模目标鼓励微调模型学习生成式表征,在三个下游NLU任务的域内和域外场景下,与多个强基线方法相比,实现了具有竞争力的准确率和最低的期望校准误差。