Automatic headline generation enables users to comprehend ongoing news events promptly and has recently become an important task in web mining and natural language processing. With the growing need for news headline generation, we argue that the hallucination issue, namely the generated headlines being not supported by the original news stories, is a critical challenge for the deployment of this feature in web-scale systems Meanwhile, due to the infrequency of hallucination cases and the requirement of careful reading for raters to reach the correct consensus, it is difficult to acquire a large dataset for training a model to detect such hallucinations through human curation. In this work, we present a new framework named ExHalder to address this challenge for headline hallucination detection. ExHalder adapts the knowledge from public natural language inference datasets into the news domain and learns to generate natural language sentences to explain the hallucination detection results. To evaluate the model performance, we carefully collect a dataset with more than six thousand labeled <article, headline> pairs. Extensive experiments on this dataset and another six public ones demonstrate that ExHalder can identify hallucinated headlines accurately and justifies its predictions with human-readable natural language explanations.
翻译:自动标题生成技术使用户能够快速理解正在发生的新闻事件,近年来已成为网络挖掘和自然语言处理领域的重要任务。随着新闻标题生成需求的日益增长,我们认为幻觉问题——即生成的标题与原始新闻内容不符——是这一功能在网络级系统中部署的关键挑战。同时,由于幻觉案例的稀缺性以及评估者需仔细阅读才能达成正确共识,通过人工标注获取大规模训练数据集来检测此类幻觉十分困难。本文提出新框架ExHalder以应对新闻标题幻觉检测的挑战。ExHalder将公开自然语言推理数据集的知识迁移至新闻领域,并学习生成自然语言语句以解释幻觉检测结果。为评估模型性能,我们精心构建了包含六千余条标注<文章,标题>对的数据集。在该数据集及六个公开数据集上的大量实验表明,ExHalder能够准确识别幻觉标题,并通过人类可读的自然语言解释验证其预测结果。