Large Language Models (LLMs) are increasingly deployed in real-world applications that require access to up-to-date knowledge. However, retraining LLMs is computationally expensive. Therefore, knowledge editing techniques are crucial for maintaining current information and correcting erroneous assertions within pre-trained models. Current benchmarks for knowledge editing primarily focus on recalling edited facts, often neglecting their logical consequences. To address this limitation, we introduce a new benchmark designed to evaluate how knowledge editing methods handle the logical consequences of a single fact edit. Our benchmark extracts relevant logical rules from a knowledge graph for a given edit. Then, it generates multi-hop questions based on these rules to assess the impact on logical consequences. Our findings indicate that while existing knowledge editing approaches can accurately insert direct assertions into LLMs, they frequently fail to inject entailed knowledge. Specifically, experiments with popular methods like ROME and FT reveal a substantial performance gap, up to 24%, between evaluations on directly edited knowledge and on entailed knowledge. This highlights the critical need for semantics-aware evaluation frameworks in knowledge editing.
翻译:大语言模型(LLMs)在需要获取最新知识的实际应用中部署日益增多。然而,重新训练LLMs的计算成本高昂。因此,知识编辑技术对于在预训练模型中维护最新信息并纠正错误断言至关重要。当前的知识编辑基准主要关注对已编辑事实的回忆,往往忽略其逻辑推论。为解决这一局限,我们引入了一个新基准,旨在评估知识编辑方法如何处理单个事实编辑的逻辑推论。我们的基准从知识图谱中提取与给定编辑相关的逻辑规则,然后基于这些规则生成多跳问题,以评估对逻辑推论的影响。我们的研究结果表明,现有知识编辑方法能够准确地将直接断言注入LLMs,但常常无法注入蕴含知识。具体而言,使用ROME和FT等流行方法进行的实验显示,直接编辑知识的评估与蕴含知识的评估之间存在高达24%的显著性能差距。这凸显了知识编辑中对语义感知评估框架的迫切需求。