Measuring the alignment between a Knowledge Graph (KG) and Large Language Models (LLMs) is an effective method to assess the factualness and identify the knowledge blind spots of LLMs. However, this approach encounters two primary challenges including the translation of KGs into natural language and the efficient evaluation of these extensive and complex structures. In this paper, we present KGLens--a novel framework aimed at measuring the alignment between KGs and LLMs, and pinpointing the LLMs' knowledge deficiencies relative to KGs. KGLens features a graph-guided question generator for converting KGs into natural language, along with a carefully designed sampling strategy based on parameterized KG structure to expedite KG traversal. We conducted experiments using three domain-specific KGs from Wikidata, which comprise over 19,000 edges, 700 relations, and 21,000 entities. Our analysis across eight LLMs reveals that KGLens not only evaluates the factual accuracy of LLMs more rapidly but also delivers in-depth analyses on topics, temporal dynamics, and relationships. Furthermore, human evaluation results indicate that KGLens can assess LLMs with a level of accuracy nearly equivalent to that of human annotators, achieving 95.7% of the accuracy rate.
翻译:衡量知识图谱(KG)与大语言模型(LLM)之间的对齐程度,是评估LLM事实准确性并识别其知识盲区的有效方法。然而,该方法面临两项主要挑战:将KG转化为自然语言,以及对这些庞大复杂结构的高效评估。本文提出KGLens——一种旨在测量KG与LLM对齐度,并定位LLM相对于KG的知识缺陷的新型框架。KGLens通过图引导的问题生成器将KG转化为自然语言,并基于参数化KG结构精心设计采样策略,以加速KG遍历。我们使用来自Wikidata的三个领域特定KG开展实验,这些KG包含超过19,000条边、700种关系及21,000个实体。对八个LLM的分析表明,KGLens不仅能更快评估LLM的事实准确性,还能提供关于主题、时间动态和关系的深度分析。此外,人工评估结果显示,KGLens评估LLM的准确率几乎与人类标注者相当,达到了95.7%的准确率。