Although achieving great success, Large Language Models (LLMs) usually suffer from unreliable hallucinations. In this paper, we define a new task of Knowledge-aware Language Model Attribution (KaLMA) that improves upon three core concerns on conventional attributed LMs. First, we extend attribution source from unstructured texts to Knowledge Graph (KG), whose rich structures benefit both the attribution performance and working scenarios. Second, we propose a new ``Conscious Incompetence" setting considering the incomplete knowledge repository, where the model identifies the need for supporting knowledge beyond the provided KG. Third, we propose a comprehensive automatic evaluation metric encompassing text quality, citation quality, and text citation alignment. To implement the above innovations, we build a dataset in biography domain BioKaLMA via a well-designed evolutionary question generation strategy, to control the question complexity and necessary knowledge to the answer. For evaluation, we develop a baseline solution and demonstrate the room for improvement in LLMs' citation generation, emphasizing the importance of incorporating the "Conscious Incompetence" setting, and the critical role of retrieval accuracy.
翻译:尽管取得了巨大成功,大语言模型(LLMs)通常存在不可靠的幻觉问题。本文定义了一项名为“知识感知语言模型归因”(KaLMA)的新任务,该任务针对传统归因语言模型的三个核心问题进行了改进。首先,我们将归因来源从非结构化文本扩展到知识图谱(KG),其丰富的结构既有利于提升归因性能,也拓展了应用场景。其次,我们提出了一种新的“有意识的无能”(Conscious Incompetence)设置,考虑了不完整知识库的情况,使模型能够识别超出给定知识图谱范围所需的支持知识。第三,我们提出了一种综合性的自动评估指标,涵盖文本质量、引文质量以及文本-引文对齐度。为实现上述创新,我们通过精心设计的演化式问题生成策略,在传记领域构建了BioKaLMA数据集,以控制问题的复杂度和回答所需的知识量。在评估方面,我们开发了一个基线解决方案,并展示了LLMs在引文生成方面仍有改进空间,强调了引入“有意识的无能”设置的重要性,以及检索准确性的关键作用。