Expert-designed close-ended benchmarks serve as vital tools in assessing the knowledge capacity of large language models (LLMs). Despite their widespread use, concerns have mounted regarding their reliability due to limited test scenarios and an unavoidable risk of data contamination. To rectify this, we present PertEval, a toolkit devised for in-depth probing of LLMs' knowledge capacity through knowledge-invariant perturbations. These perturbations employ human-like restatement techniques to generate on-the-fly test samples from static benchmarks, meticulously retaining knowledge-critical content while altering irrelevant details. Our toolkit further includes a suite of transition analyses that compare performance on raw vs. perturbed test sets to precisely assess LLMs' genuine knowledge capacity. Six state-of-the-art LLMs are re-evaluated using PertEval. Results reveal significantly inflated performance of the LLMs on raw benchmarks, including an absolute 21% overestimation for GPT-4. Additionally, through a nuanced response pattern analysis, we discover that PertEval retains LLMs' uncertainty to specious knowledge, potentially being resolved through rote memorization and leading to inflated performance. We also find that the detailed transition analyses by PertEval could illuminate weaknesses in existing LLMs' knowledge mastery and guide the development of refinement. Given these insights, we posit that PertEval can act as an essential tool that, when applied alongside any close-ended benchmark, unveils the true knowledge capacity of LLMs, marking a significant step toward more trustworthy LLM evaluation.
翻译:专家设计的封闭式基准测试是评估大型语言模型知识能力的重要工具。尽管其应用广泛,但由于测试场景有限及不可避免的数据污染风险,其可靠性日益受到质疑。为此,我们提出了PertEval工具包,旨在通过知识不变扰动深入探查大型语言模型的知识能力。该扰动技术采用类人类重述方法,从静态基准动态生成测试样本,在严格保留知识核心内容的同时改变无关细节。本工具包还包含一套转移分析组件,通过对比原始测试集与扰动测试集的性能表现,精准评估大型语言模型的真实知识能力。使用PertEval对六个前沿大型语言模型进行重新评估的结果显示,这些模型在原始基准上的性能被显著高估,其中GPT-4的绝对高估幅度达21%。此外,通过细粒度的响应模式分析,我们发现PertEval能有效保留大型语言模型对虚假知识的不确定性——这种不确定性可能通过机械记忆被消除,从而导致性能虚高。研究还表明,PertEval提供的精细化转移分析能够揭示现有大型语言模型知识掌握中的薄弱环节,并为模型优化提供指导。基于这些发现,我们认为PertEval可成为关键评估工具:当与任意封闭式基准测试结合使用时,能够揭示大型语言模型的真实知识能力,这标志着朝着更可信的大型语言模型评估迈出了重要一步。