Recently proposed long-form question answering (QA) systems, supported by large language models (LLMs), have shown promising capabilities. Yet, attributing and verifying their generated abstractive answers can be difficult, and automatically evaluating their accuracy remains an ongoing challenge. In this work, we introduce a new QA task for answering multi-answer questions by summarizing multiple diverse sources in a semi-extractive fashion. Specifically, Semi-extractive Multi-source QA (SEMQA) requires models to output a comprehensive answer, while mixing factual quoted spans -- copied verbatim from given input sources -- and non-factual free-text connectors that glue these spans together into a single cohesive passage. This setting bridges the gap between the outputs of well-grounded but constrained extractive QA systems and more fluent but harder to attribute fully abstractive answers. Particularly, it enables a new mode for language models that leverages their advanced language generation capabilities, while also producing fine in-line attributions by-design that are easy to verify, interpret, and evaluate. To study this task, we create the first dataset of this kind, QuoteSum, with human-written semi-extractive answers to natural and generated questions, and define text-based evaluation metrics. Experimenting with several LLMs in various settings, we find this task to be surprisingly challenging, demonstrating the importance of QuoteSum for developing and studying such consolidation capabilities.
翻译:近期提出的基于大语言模型(LLM)的长文本问答系统已展现出良好的性能。然而,对其生成的抽象式答案进行归因和验证较为困难,且自动评估其准确性仍是一个持续的挑战。本研究针对多答案问题提出一种新的问答任务,即通过半抽取式方法对多个多样化来源进行摘要。具体而言,半抽取式多源问答要求模型输出综合性答案,同时混合事实性引用片段(从给定输入源逐字复制)与非事实性自由文本连接词,将这些片段整合为连贯的单一文本。该设定在输出受限但基础扎实的抽取式问答系统与流畅但难以完全归因的抽象式答案之间架起了桥梁。特别地,它为大语言模型提供了一种新模式:既能发挥其先进的语言生成能力,又能通过设计产生易于验证、解释和评估的精细行内归因。为研究该任务,我们创建了首个此类数据集QuoteSum,其中包含针对自然问题与生成问题的人工撰写半抽取式答案,并定义了基于文本的评估指标。通过对多种设置下的多个大语言模型进行实验,我们发现该任务具有出乎意料的挑战性,这证明了QuoteSum对于开发和研究此类信息整合能力的重要性。