As language models are adapted by a more sophisticated and diverse set of users, the importance of guaranteeing that they provide factually correct information supported by verifiable sources is critical across fields of study & professions. This is especially the case for high-stakes fields, such as medicine and law, where the risk of propagating false information is high and can lead to undesirable societal consequences. Previous work studying factuality and attribution has not focused on analyzing these characteristics of language model outputs in domain-specific scenarios. In this work, we present an evaluation study analyzing various axes of factuality and attribution provided in responses from a few systems, by bringing domain experts in the loop. Specifically, we first collect expert-curated questions from 484 participants across 32 fields of study, and then ask the same experts to evaluate generated responses to their own questions. We also ask experts to revise answers produced by language models, which leads to ExpertQA, a high-quality long-form QA dataset with 2177 questions spanning 32 fields, along with verified answers and attributions for claims in the answers.
翻译:随着语言模型被更复杂和多样化的用户群体所采用,确保它们提供有可验证来源支持的事实正确信息,在各学科领域及专业中变得至关重要。这在医学和法律等高风险领域尤为突出,因为在这些领域传播错误信息的风险很高,且可能导致不良社会后果。先前关于事实性和归属性的研究并未专注于分析语言模型输出在特定领域场景中的这些特性。在本工作中,我们通过引入领域专家参与,提出了一项评估研究,分析多个系统在响应中提供的事实性和归属性的多个维度。具体而言,我们首先从涵盖32个研究领域的484名参与者收集专家策展的问题,然后要求同一批专家评估对其自身问题生成的响应。我们还要求专家修订语言模型生成的答案,从而产生了ExpertQA——一个高质量的长形式问答数据集,包含涵盖32个领域的2177个问题,以及答案中声明的验证答案和归属性信息。