Knowledge graph embedding models (KGEMs) have gained considerable traction in recent years. These models learn a vector representation of knowledge graph entities and relations, a.k.a. knowledge graph embeddings (KGEs). Learning versatile KGEs is desirable as it makes them useful for a broad range of tasks. However, KGEMs are usually trained for a specific task, which makes their embeddings task-dependent. In parallel, the widespread assumption that KGEMs actually create a semantic representation of the underlying entities and relations (e.g., project similar entities closer than dissimilar ones) has been challenged. In this work, we design heuristics for generating protographs -- small, modified versions of a KG that leverage RDF/S information. The learnt protograph-based embeddings are meant to encapsulate the semantics of a KG, and can be leveraged in learning KGEs that, in turn, also better capture semantics. Extensive experiments on various evaluation benchmarks demonstrate the soundness of this approach, which we call Modular and Agnostic SCHema-based Integration of protograph Embeddings (MASCHInE). In particular, MASCHInE helps produce more versatile KGEs that yield substantially better performance for entity clustering and node classification tasks. For link prediction, using MASCHinE substantially increases the number of semantically valid predictions with equivalent rank-based performance.
翻译:知识图谱嵌入模型(KGEMs)近年来获得了显著关注。这些模型学习知识图谱实体和关系的向量表示,即知识图谱嵌入(KGEs)。学习通用KGEs具有重要价值,因为它能使嵌入广泛应用于多种任务。然而,KGEMs通常针对特定任务进行训练,导致其嵌入具有任务依赖性。与此同时,一个普遍假设——即KGEMs实际上能为底层实体和关系创建语义表示(例如,将相似实体投影到比不相似实体更近的位置)——已受到质疑。本文设计了生成原型图(protographs)的启发式方法——这些原型图是基于RDF/S信息对知识图谱进行小规模修改后的版本。基于原型图学习的嵌入旨在封装知识图谱的语义,并可被用于学习能更好捕获语义的KGEs。在多种评估基准上的大量实验证明了该方法(我们将其命名为模块化与无关的基于语义模式的原型图嵌入集成方法,MASCHInE)的合理性。特别地,MASCHInE有助于生成更通用的KGEs,在实体聚类和节点分类任务上取得显著更优的性能。对于链接预测任务,使用MASCHInE在保持等价排序性能的同时,大幅增加了语义有效预测的数量。