Abstractive summarization aims at generating natural language summaries of a source document that are succinct while preserving the important elements. Despite recent advances, neural text summarization models are known to be susceptible to hallucinating (or more correctly confabulating), that is to produce summaries with details that are not grounded in the source document. In this paper, we introduce a simple yet efficient technique, CoBa, to reduce hallucination in abstractive summarization. The approach is based on two steps: hallucination detection and mitigation. We show that the former can be achieved through measuring simple statistics about conditional word probabilities and distance to context words. Further, we demonstrate that straight-forward backtracking is surprisingly effective at mitigation. We thoroughly evaluate the proposed method with prior art on three benchmark datasets for text summarization. The results show that CoBa is effective and efficient in reducing hallucination, and offers great adaptability and flexibility.
翻译:抽象摘要旨在生成源文档的自然语言摘要,在保留重要信息的同时保持简洁。尽管近来取得了进展,但神经文本摘要模型仍易产生幻觉(更准确地说,是虚构现象),即生成的摘要中包含源文档中未提及的细节。本文提出一种简单高效的CoBa技术,以减少抽象摘要中的幻觉生成。该方法基于两个步骤:幻觉检测与缓解。研究表明,前者可通过测量条件词概率和上下文词距离的简单统计量实现。此外,我们发现直接回溯在缓解幻觉方面出奇有效。我们在三个文本摘要基准数据集上,将所提方法与现有技术进行全面评估。结果表明,CoBa在减少幻觉方面既有效又高效,并展现出极佳的适应性与灵活性。