We present a knowledge integration framework (called KIF) that uses Wikidata as a lingua franca to integrate heterogeneous knowledge bases. These can be triplestores, relational databases, CSV files, etc., which may or may not use the Wikidata dialect of RDF. KIF leverages Wikidata's data model and vocabulary plus user-defined mappings to expose a unified view of the integrated bases while keeping track of the context and provenance of their statements. The result is a virtual knowledge base which behaves like an "extended Wikidata" and which can be queried either through an efficient filter interface or using SPARQL. We present the design and implementation of KIF, discuss how we have used it to solve a real integration problem in the domain of chemistry (involving Wikidata, PubChem, and IBM CIRCA), and present experimental results on the performance and overhead of KIF.
翻译:我们提出了一种知识集成框架(称为KIF),该框架以Wikidata作为通用语言来集成异构知识库。这些知识库可以是三元组存储、关系数据库、CSV文件等,它们可能使用也可能未使用Wikidata的RDF方言。KIF利用Wikidata的数据模型和词汇表,结合用户定义的映射规则,在追踪各陈述语境与出处信息的同时,对集成后的知识库呈现统一视图。其成果是一个行为类似于"扩展版Wikidata"的虚拟知识库,既可通过高效的筛选接口查询,也可通过SPARQL进行检索。本文介绍了KIF的设计与实现,阐述了如何将其用于解决化学领域真实集成问题(涉及Wikidata、PubChem和IBM CIRCA),并给出了关于KIF性能与开销的实验结果。