While the NLP community has produced numerous summarization benchmarks, none provide the rich annotations required to simultaneously address many important problems related to control and reliability. We introduce a Wikipedia-derived benchmark, complemented by a rich set of crowd-sourced annotations, that supports $8$ interrelated tasks: (i) extractive summarization; (ii) abstractive summarization; (iii) topic-based summarization; (iv) compressing selected sentences into a one-line summary; (v) surfacing evidence for a summary sentence; (vi) predicting the factual accuracy of a summary sentence; (vii) identifying unsubstantiated spans in a summary sentence; (viii) correcting factual errors in summaries. We compare various methods on this benchmark and discover that on multiple tasks, moderately-sized fine-tuned models consistently outperform much larger few-shot prompted language models. For factuality-related tasks, we also evaluate existing heuristics to create training data and find that training on them results in worse performance than training on $20\times$ less human-labeled data. Our articles draw from $6$ domains, facilitating cross-domain analysis. On some tasks, the amount of training data matters more than the domain where it comes from, while for other tasks training specifically on data from the target domain, even if limited, is more beneficial.
翻译:尽管自然语言处理社区已产出大量摘要基准,但尚无基准能提供同时解决控制与可靠性相关诸多重要问题所需的丰富注释。我们引入一个基于维基百科的基准,辅以丰富的众包注释,支持8项相互关联的任务:(i) 抽取式摘要;(ii) 生成式摘要;(iii) 基于主题的摘要;(iv) 将选定句子压缩为一行摘要;(v) 为摘要句提供证据;(vi) 预测摘要句的事实准确性;(vii) 识别摘要句中的无依据片段;(viii) 纠正摘要中的事实错误。我们在此基准上比较多种方法,发现在多项任务中,中等规模的微调模型始终优于规模大得多的少样本提示语言模型。针对事实性相关任务,我们评估了现有用于创建训练数据的启发式方法,发现基于这些数据训练的效果,甚至不如使用20倍更少的人工标注数据。文章涵盖6个领域,便于跨领域分析。在某些任务中,训练数据量比其来源领域更为重要;而在其他任务中,即使是有限的特定领域训练数据,也能带来更优的效果。