There has been a rapid growth in biomedical literature, yet capturing the heterogeneity of the bibliographic information of these articles remains relatively understudied. Although graph mining research via heterogeneous graph neural networks has taken center stage, it remains unclear whether these approaches capture the heterogeneity of the PubMed database, a vast digital repository containing over 33 million articles. We introduce PubMed Graph Benchmark (PGB), a new benchmark dataset for evaluating heterogeneous graph embeddings for biomedical literature. PGB is one of the largest heterogeneous networks to date and consists of 30 million English articles. The benchmark contains rich metadata including abstract, authors, citations, MeSH terms, MeSH hierarchy, and some other information. The benchmark contains an evaluation task of 21 systematic reviews topics from 3 different datasets. In PGB, we aggregate the metadata associated with the biomedical articles from PubMed into a unified source and make the benchmark publicly available for any future works.
翻译:摘要:生物医学文献数量快速增长,但如何捕捉这些文章书目信息的异质性仍是一个相对研究不足的问题。尽管基于异构图神经网络的图挖掘研究已成为焦点,但这些方法能否有效捕捉PubMed数据库(包含超过3300万篇文章的海量数字存储库)的异质性仍不明确。我们提出了PubMed图基准(PGB),这是一个用于评估生物医学文献异构图嵌入的新基准数据集。PGB是迄今为止规模最大的异构网络之一,包含3000万篇英文文章。该基准包含丰富的元数据,包括摘要、作者、引用、MeSH术语、MeSH层级结构及其他信息。基准测试任务涵盖来自3个不同数据集的21个系统综述主题。在PGB中,我们将PubMed生物医学文章相关的元数据整合至统一数据源,并公开该基准供未来研究使用。