In recent years, we have witnessed the proliferation of knowledge graphs (KG) in various domains, aiming to support applications like question answering, recommendations, etc. A frequent task when integrating knowledge from different KGs is to find which subgraphs refer to the same real-world entity. Recently, embedding methods have been used for entity alignment tasks, that learn a vector-space representation of entities which preserves their similarity in the original KGs. A wide variety of supervised, unsupervised, and semi-supervised methods have been proposed that exploit both factual (attribute based) and structural information (relation based) of entities in the KGs. Still, a quantitative assessment of their strengths and weaknesses in real-world KGs according to different performance metrics and KG characteristics is missing from the literature. In this work, we conduct the first meta-level analysis of popular embedding methods for entity alignment, based on a statistically sound methodology. Our analysis reveals statistically significant correlations of different embedding methods with various meta-features extracted by KGs and rank them in a statistically significant way according to their effectiveness across all real-world KGs of our testbed. Finally, we study interesting trade-offs in terms of methods' effectiveness and efficiency.
翻译:近年来,知识图谱在多个领域蓬勃发展,旨在支持问答、推荐等应用。整合不同知识图谱时,一个常见任务是判断哪些子图指向同一真实世界实体。近期,嵌入方法被用于实体对齐任务,通过学习实体的向量空间表示,保留其在原始知识图谱中的相似性。目前已有多种监督、无监督及半监督方法被提出,这些方法同时利用了知识图谱中实体的事实(基于属性)和结构(基于关系)信息。然而,现有文献缺乏根据不同性能指标和知识图谱特征对实际知识图谱中这些方法优劣的定量评估。本研究基于统计严谨的方法,首次对主流实体对齐嵌入方法进行元层次分析。我们的分析揭示了不同嵌入方法与从知识图谱中提取的各类元特征之间存在统计显著的关联,并根据所有测试知识图谱上的有效性,以统计显著的方式对这些方法进行了排序。最后,我们探讨了方法在有效性和效率之间的有趣权衡。