Despite their impressive scale, knowledge bases (KBs), such as Wikidata, still contain significant gaps. Language models (LMs) have been proposed as a source for filling these gaps. However, prior works have focused on prominent entities with rich coverage by LMs, neglecting the crucial case of long-tail entities. In this paper, we present a novel method for LM-based-KB completion that is specifically geared for facts about long-tail entities. The method leverages two different LMs in two stages: for candidate retrieval and for candidate verification and disambiguation. To evaluate our method and various baselines, we introduce a novel dataset, called MALT, rooted in Wikidata. Our method outperforms all baselines in F1, with major gains especially in recall.
翻译:尽管知识库(如Wikidata)在规模上令人瞩目,但仍存在显著的信息空缺。语言模型(LM)已被提出可用于填补这些空缺。然而,先前的研究主要聚焦于被语言模型丰富覆盖的显著实体,忽略了关键的长尾实体情形。本文提出了一种面向语言模型的知识库补全新方法,专门针对长尾实体的知识事实。该方法分两阶段利用两种不同的语言模型:候选检索阶段与候选验证及消歧阶段。为评估本方法与各类基线模型,我们构建了一个基于Wikidata的新数据集MALT。实验表明,本方法在F1分数上优于所有基线方法,尤其在召回率方面取得了显著提升。