Despite recent progress in abstractive summarization, models often generate summaries with factual errors. Numerous approaches to detect these errors have been proposed, the most popular of which are question answering (QA)-based factuality metrics. These have been shown to work well at predicting summary-level factuality and have potential to localize errors within summaries, but this latter capability has not been systematically evaluated in past research. In this paper, we conduct the first such analysis and find that, contrary to our expectations, QA-based frameworks fail to correctly identify error spans in generated summaries and are outperformed by trivial exact match baselines. Our analysis reveals a major reason for such poor localization: questions generated by the QG module often inherit errors from non-factual summaries which are then propagated further into downstream modules. Moreover, even human-in-the-loop question generation cannot easily offset these problems. Our experiments conclusively show that there exist fundamental issues with localization using the QA framework which cannot be fixed solely by stronger QA and QG models.
翻译:尽管抽象式摘要生成近期取得进展,模型生成的摘要仍常包含事实性错误。已有多种检测此类错误的方法被提出,其中最流行的是基于问答的事实性度量指标。研究表明,这些指标在预测摘要整体事实性方面效果良好,并具有定位摘要内部错误位置的潜力,但这一能力尚未在以往研究中进行系统评估。本文首次开展此类分析,发现与预期相反,基于问答的框架无法正确识别生成摘要中的错误片段,其表现甚至不如简单的精确匹配基线方法。进一步分析揭示其定位能力低下的主要原因:问答生成模块产生的问题往往继承自存在事实错误的摘要,并进一步传播至下游模块。此外,即使采用人在回路式的问题生成方式,也难以轻易弥补这些问题。我们的实验最终证实,基于问答框架的错误定位存在本质性缺陷,单纯依赖更强的问答与问答生成模型无法解决该问题。