One of the significant barriers to the training of statistical models on knowledge graphs is the difficulty that scientists have in finding the best input data to address their prediction goal. In addition to this, a key challenge is to determine how to manipulate these relational data, which are often in the form of particular triples (i.e., subject, predicate, object), to enable the learning process. Currently, many high-quality catalogs of knowledge graphs, are available. However, their primary goal is the re-usability of these resources, and their interconnection, in the context of the Semantic Web. This paper describes the LiveSchema initiative, namely, a first version of a gateway that has the main scope of leveraging the gold mine of data collected by many existing catalogs collecting relational data like ontologies and knowledge graphs. At the current state, LiveSchema contains - 1000 datasets from 4 main sources and offers some key facilities, which allow to: i) evolving LiveSchema, by aggregating other source catalogs and repositories as input sources; ii) querying all the collected resources; iii) transforming each given dataset into formal concept analysis matrices that enable analysis and visualization services; iv) generating models and tensors from each given dataset.
翻译:训练知识图谱统计模型的主要障碍之一在于科学家难以找到最适合预测目标的最佳输入数据。此外,如何操控这些通常以特定三元组(即主语、谓语、宾语)形式呈现的关系数据以支持学习过程,也是一项关键挑战。目前虽存在许多高质量的知识图谱目录,但其主要目标是在语义网背景下实现资源的可重用性与互联互通。本文介绍了LiveSchema项目——一个初步版本的网关,其核心目标是利用现有众多目录收集的关系数据(如本体与知识图谱)这座数据金矿。当前版本中,LiveSchema包含来自4个主要来源的1000个数据集,并提供以下关键功能:i)通过聚合其他来源目录与存储库作为输入源,持续演进LiveSchema;ii)对所有收集的资源进行查询;iii)将每个给定数据集转换为形式概念分析矩阵,支持分析与可视化服务;iv)从每个给定数据集生成模型与张量。