Abstractive summarization models often generate factually inconsistent content particularly when the parametric knowledge of the model conflicts with the knowledge in the input document. In this paper, we analyze the robustness of fine-tuning based summarization models to the knowledge conflict, which we call factual adaptiveness. We utilize pre-trained language models to construct evaluation sets and find that factual adaptiveness is not strongly correlated with factual consistency on original datasets. Furthermore, we introduce a controllable counterfactual data augmentation method where the degree of knowledge conflict within the augmented data can be adjustable. Our experimental results on two pre-trained language models (PEGASUS and BART) and two fine-tuning datasets (XSum and CNN/DailyMail) demonstrate that our method enhances factual adaptiveness while achieving factual consistency on original datasets on par with the contrastive learning baseline.
翻译:抽象式摘要模型常生成与事实不一致的内容,尤其在模型参数知识与输入文档知识冲突时。本文分析了基于微调的摘要模型对知识冲突的鲁棒性,我们称之为事实适应性。利用预训练语言模型构建评估集,我们发现事实适应性与原始数据集上的事实一致性并无强关联。此外,我们引入一种可控的反事实数据增强方法,其中增强数据中的知识冲突程度可调节。在两种预训练语言模型(PEGASUS和BART)及两个微调数据集(XSum和CNN/DailyMail)上的实验结果表明,我们的方法在提升事实适应性的同时,实现了与对比学习基线相当的事实一致性。