Research publications are the primary vehicle for sharing scientific progress in the form of new discoveries, methods, techniques, and insights. Unfortunately, the lack of a large-scale, comprehensive, and easy-to-use resource capturing the myriad relationships between publications, their authors, and venues presents a barrier to applications for gaining a deeper understanding of science. In this paper, we present PubGraph, a new resource for studying scientific progress that takes the form of a large-scale knowledge graph (KG) with more than 385M entities, 13B main edges, and 1.5B qualifier edges. PubGraph is comprehensive and unifies data from various sources, including Wikidata, OpenAlex, and Semantic Scholar, using the Wikidata ontology. Beyond the metadata available from these sources, PubGraph includes outputs from auxiliary community detection algorithms and large language models. To further support studies on reasoning over scientific networks, we create several large-scale benchmarks extracted from PubGraph for the core task of knowledge graph completion (KGC). These benchmarks present many challenges for knowledge graph embedding models, including an adversarial community-based KGC evaluation setting, zero-shot inductive learning, and large-scale learning. All of the aforementioned resources are accessible at https://pubgraph.isi.edu/ and released under the CC-BY-SA license. We plan to update PubGraph quarterly to accommodate the release of new publications.
翻译:科研论文是分享新发现、方法、技术与见解等科学进展的主要载体。然而,缺乏大规模、全面且易于使用的资源来捕捉论文、作者及会议期刊之间的多重关系,这为深入理解科学的应用造成了障碍。本文提出PubGraph——一种研究科学进展的新资源,它以大规模知识图谱(KG)的形式呈现,包含超过3.85亿个实体、130亿条主边和15亿条限定符边。PubGraph全面整合了来自Wikidata、OpenAlex和Semantic Scholar等多个来源的数据,并采用Wikidata本体。除了这些来源提供的元数据外,PubGraph还囊括了辅助社区检测算法和大语言模型的输出结果。为进一步支持对科学网络推理的研究,我们基于PubGraph创建了多个大规模基准测试,用于知识图谱补全(KGC)这一核心任务。这些基准测试对知识图谱嵌入模型提出了诸多挑战,包括对抗性基于社区的KGC评估设置、零样本归纳学习以及大规模学习。上述所有资源均可通过https://pubgraph.isi.edu/ 获取,并依据CC-BY-SA许可协议发布。我们计划按季度更新PubGraph,以适配新出版物的发布。