Fine-tuning large language models can improve task specific performance, although a general understanding of what the fine-tuned model has learned, forgotten and how to trust its predictions is still missing. We derive principled uncertainty quantification for fine-tuned LLMs with posterior approximations using computationally efficient low-rank adaptation ensembles. We analyze three common multiple-choice datasets using low-rank adaptation ensembles based on Mistral-7b, and draw quantitative and qualitative conclusions on their perceived complexity and model efficacy on the different target domains during and after fine-tuning. In particular, backed by the numerical experiments, we hypothesise about signals from entropic uncertainty measures for data domains that are inherently difficult for a given architecture to learn.
翻译:微调大型语言模型可以提升特定任务的性能,但关于微调后的模型学到了什么、遗忘了什么以及如何信任其预测,仍缺乏系统性理解。我们通过利用计算高效的低秩适应集成方法,基于后验近似推导出微调大语言模型的严谨不确定性量化。基于Mistral-7b的低秩适应集成,我们分析了三个常见的多项选择数据集,并在微调过程中和微调后,定量与定性评估了这些数据集在不同目标领域中的感知复杂性与模型效能。特别地,基于数值实验,我们提出假设:对于给定架构而言,数据域中某些固有难以学习的部分,可以通过熵不确定性度量获取其信号。