Applications of large open-domain knowledge graphs (KGs) to real-world problems pose many unique challenges. In this paper, we present extensions to Saga our platform for continuous construction and serving of knowledge at scale. In particular, we describe a pipeline for training knowledge graph embeddings that powers key capabilities such as fact ranking, fact verification, a related entities service, and support for entity linking. We then describe how our platform, including graph embeddings, can be leveraged to create a Semantic Annotation service that links unstructured Web documents to entities in our KG. Semantic annotation of the Web effectively expands our knowledge graph with edges to open-domain Web content which can be used in various search and ranking problems. Finally, we leverage annotated Web documents to drive Open-domain Knowledge Extraction. This targeted extraction framework identifies important coverage issues in the KG, then finds relevant data sources for target entities on the Web and extracts missing information to enrich the KG. Finally, we describe adaptations to our knowledge platform needed to construct and serve private personal knowledge on-device. This includes private incremental KG construction, cross-device knowledge sync, and global knowledge enrichment.
翻译:大型开放域知识图谱(KG)在解决现实世界问题中的应用带来了诸多独特挑战。本文介绍了我们对Saga平台(用于大规模持续构建与服务知识的平台)的扩展。具体而言,我们描述了一个训练知识图谱嵌入的流水线,该流水线支持事实排序、事实验证、相关实体服务以及实体链接等关键功能。随后,我们阐述了如何利用该平台(包括图嵌入)构建语义标注服务,将非结构化Web文档链接到知识图谱中的实体。对Web的语义标注通过连接至开放域Web内容的边有效扩展了知识图谱,这些边可用于多种搜索与排序问题。此外,我们利用标注后的Web文档驱动开放域知识抽取。这一目标导向的抽取框架识别知识图谱中的重要覆盖缺口,随后在Web上为目标实体寻找相关数据源,并抽取缺失信息以丰富知识图谱。最后,我们描述了为在设备端构建与服务私有个人知识所需的知识平台适配方案,包括私有增量知识图谱构建、跨设备知识同步及全局知识丰富。