Knowledge graphs represent real-world entities and their relations in a semantically-rich structure supported by ontologies. Exploring this data with machine learning methods often relies on knowledge graph embeddings, which produce latent representations of entities that preserve structural and local graph neighbourhood properties, but sacrifice explainability. However, in tasks such as link or relation prediction, understanding which specific features better explain a relation is crucial to support complex or critical applications. We propose SEEK, a novel approach for explainable representations to support relation prediction in knowledge graphs. It is based on identifying relevant shared semantic aspects (i.e., subgraphs) between entities and learning representations for each subgraph, producing a multi-faceted and explainable representation. We evaluate SEEK on two real-world highly complex relation prediction tasks: protein-protein interaction prediction and gene-disease association prediction. Our extensive analysis using established benchmarks demonstrates that SEEK achieves significantly better performance than standard learning representation methods while identifying both sufficient and necessary explanations based on shared semantic aspects.
翻译:知识图谱通过本体论支持的语义丰富结构,表示现实世界实体及其关系。利用机器学习方法探索此类数据时,通常依赖知识图谱嵌入技术——该技术可生成保留结构特性及局部图邻域属性的实体潜在表示,但牺牲了可解释性。然而,在链接预测或关系预测等任务中,理解哪些特定特征能更好解释某种关系,对于支撑复杂或关键性应用至关重要。我们提出SEEK方法,一种面向知识图谱关系预测的新型可解释表示方法。该方法通过识别实体间相关的共享语义方面(即子图),并为每个子图学习相应表示,从而生成多层面且可解释的表示。我们在两个真实世界的高度复杂关系预测任务上评估SEEK:蛋白质-蛋白质相互作用预测与基因-疾病关联预测。基于权威基准的广泛分析表明,SEEK不仅比标准学习表示方法取得显著更优性能,还能基于共享语义方面识别充分且必要的解释。