Previous studies have relied on existing question-answering benchmarks to evaluate the knowledge stored in large language models (LLMs). However, this approach has limitations regarding factual knowledge coverage, as it mostly focuses on generic domains which may overlap with the pretraining data. This paper proposes a framework to systematically assess the factual knowledge of LLMs by leveraging knowledge graphs (KGs). Our framework automatically generates a set of questions and expected answers from the facts stored in a given KG, and then evaluates the accuracy of LLMs in answering these questions. We systematically evaluate the state-of-the-art LLMs with KGs in generic and specific domains. The experiment shows that ChatGPT is consistently the top performer across all domains. We also find that LLMs performance depends on the instruction finetuning, domain and question complexity and is prone to adversarial context.
翻译:先前研究依赖现有问答基准来评估大型语言模型(LLMs)中存储的知识。然而,这种方法在事实知识覆盖率方面存在局限性,因为其主要聚焦于可能与预训练数据重叠的通用领域。本文提出一个框架,通过利用知识图谱(KGs)系统性地评估LLMs的事实知识。我们的框架自动从特定知识图谱存储的事实中生成一组问题及预期答案,然后评估LLMs回答这些问题的准确性。我们使用通用领域和特定领域的知识图谱系统性评估了最先进的LLMs。实验表明,ChatGPT在所有领域中均保持最佳表现。我们还发现LLMs的性能依赖于指令微调、领域和问题复杂度,并且易受对抗性上下文的影响。