Large language models are prone to hallucinating factually incorrect statements. A key source of these errors is exposure to new factual information through supervised fine-tuning (SFT), which can increase hallucinations w.r.t. knowledge acquired during pre-training. In this work, we explore whether SFT-induced hallucinations can be mitigated using established tools from the continual learning literature, since they arise as a by-product of knowledge degradation during training. We propose a self-distillation-based SFT method that facilitates effective factual learning while minimizing hallucinations w.r.t. pre-existing knowledge by regularizing output-distribution drift. We also show that, in settings where new knowledge acquisition is unnecessary, suppressing factual plasticity by freezing parameter groups, can preserve task performance while reducing hallucinations. Lastly, we investigate the mechanism behind SFT-induced hallucinations through three hypotheses: capacity limitations, behavior cloning, and localized interference. Our experiments show that a main driver is interference among overlapping semantic representations, and that self-distillation succeeds by mitigating this interference.
翻译:大语言模型容易生成与事实不符的错误陈述。一个关键的错误来源是通过监督式微调接触新的事实信息,这可能会增加模型在预训练阶段获取知识方面的幻觉。本文探讨能否利用持续学习文献中的成熟工具来缓解监督式微调引发的幻觉,因为这种幻觉本质上是训练过程中知识退化产生的副产物。我们提出一种基于自蒸馏的监督式微调方法,通过正则化输出分布漂移,在促进有效事实学习的同时,最大程度减少对已有知识的幻觉。研究还表明,在无需获取新知识的场景中,通过冻结参数组来抑制事实可塑性,可在降低幻觉的同时保持任务性能。最后,我们从容量限制、行为克隆和局部干扰三个假设出发,探究监督式微调引发幻觉的机制。实验表明,主要驱动因素在于重叠语义表征间的相互干扰,而自蒸馏正是通过缓解这种干扰来发挥作用的。