Since the recent prosperity of Large Language Models (LLMs), there have been interleaved discussions regarding how to reduce hallucinations from LLM responses, how to increase the factuality of LLMs, and whether Knowledge Graphs (KGs), which store the world knowledge in a symbolic form, will be replaced with LLMs. In this paper, we try to answer these questions from a new angle: How knowledgeable are LLMs? To answer this question, we constructed Head-to-Tail, a benchmark that consists of 18K question-answer (QA) pairs regarding head, torso, and tail facts in terms of popularity. We designed an automated evaluation method and a set of metrics that closely approximate the knowledge an LLM confidently internalizes. Through a comprehensive evaluation of 14 publicly available LLMs, we show that existing LLMs are still far from being perfect in terms of their grasp of factual knowledge, especially for facts of torso-to-tail entities.
翻译:自大语言模型(LLM)近期蓬勃发展以来,关于如何减少LLM响应中的幻觉现象、如何提升LLM的事实准确性,以及以符号形式存储世界知识的知识图谱(KG)是否会被LLM取代等议题,始终交织在学术讨论中。本文尝试从一个全新角度回答这些问题:LLM的知识水平究竟如何?为此,我们构建了"头对尾"(Head-to-Tail)基准测试集,包含1.8万组关于头部、躯干和尾部事实(按流行度分类)的问答对。我们设计了自动化评估方法及一组度量指标,能够近似衡量LLM真正内化的知识量。通过对14个公开可用的LLM进行全面评估,我们发现现有LLM在事实知识的掌握方面仍远非完美,尤其对躯干至尾部实体相关事实的掌握存在显著不足。