Introduction: The scientific publishing landscape is expanding rapidly, creating challenges for researchers to stay up-to-date with the evolution of the literature. Natural Language Processing (NLP) has emerged as a potent approach to automating knowledge extraction from this vast amount of publications and preprints. Tasks such as Named-Entity Recognition (NER) and Named-Entity Linking (NEL), in conjunction with context-dependent semantic interpretation, offer promising and complementary approaches to extracting structured information and revealing key concepts. Results: We present the SourceData-NLP dataset produced through the routine curation of papers during the publication process. A unique feature of this dataset is its emphasis on the annotation of bioentities in figure legends. We annotate eight classes of biomedical entities (small molecules, gene products, subcellular components, cell lines, cell types, tissues, organisms, and diseases), their role in the experimental design, and the nature of the experimental method as an additional class. SourceData-NLP contains more than 620,000 annotated biomedical entities, curated from 18,689 figures in 3,223 papers in molecular and cell biology. We illustrate the dataset's usefulness by assessing BioLinkBERT and PubmedBERT, two transformers-based models, fine-tuned on the SourceData-NLP dataset for NER. We also introduce a novel context-dependent semantic task that infers whether an entity is the target of a controlled intervention or the object of measurement. Conclusions: SourceData-NLP's scale highlights the value of integrating curation into publishing. Models trained with SourceData-NLP will furthermore enable the development of tools able to extract causal hypotheses from the literature and assemble them into knowledge graphs.
翻译:引言:科学出版领域正在快速扩张,给研究人员跟踪文献发展动态带来了挑战。自然语言处理(NLP)已成为从海量出版物和预印本中自动提取知识的有效方法。命名实体识别(NER)和命名实体链接(NEL)等任务,结合上下文相关的语义解释,为提取结构化信息并揭示核心概念提供了互补且极具前景的途径。结果:我们展示了在出版流程中通过常规文献整理构建的SourceData-NLP数据集。该数据集的独特之处在于其重点标注了图表图例中的生物实体。我们标注了八类生物医学实体(小分子、基因产物、亚细胞组分、细胞系、细胞类型、组织、生物体和疾病),它们在实验设计中的作用,并将实验方法的性质作为额外类别。SourceData-NLP包含超过62万个标注的生物医学实体,这些实体来自分子和细胞生物学领域3,223篇论文中的18,689幅图表。通过评估在SourceData-NLP数据集上针对NER微调的两种基于Transformer的模型(BioLinkBERT和PubmedBERT),我们展示了该数据集的实用性。我们还引入了一种新颖的上下文相关语义任务,用于推断实体是受控干预的目标还是测量对象。结论:SourceData-NLP的规模突显了将文献整理融入出版过程的价值。使用SourceData-NLP训练的模型将进一步推动开发能够从文献中提取因果假设并将其整合到知识图谱中的工具。