Literature is the primary expression of scientific knowledge and an important source of research data. However, scientific knowledge expressed in narrative text documents is not inherently machine reusable. To facilitate knowledge reuse, e.g. for synthesis research, scientific knowledge must be extracted from articles and organized into databases post-publication. The high time costs and inaccuracies associated with completing these activities manually has driven the development of techniques that automate knowledge extraction. Tackling the problem with a different mindset, we propose a pre-publication approach, known as reborn, that ensures scientific knowledge is born reusable, i.e. produced in a machine-reusable format during knowledge production. We implement the approach using the Open Research Knowledge Graph infrastructure for FAIR scientific knowledge organization. We test the approach with three use cases, and discuss the role of publishers and editors in scaling the approach. Our results suggest that the proposed approach is superior compared to classical manual and semi-automated post-publication extraction techniques in terms of knowledge richness and accuracy as well as technological simplicity.
翻译:文献是科学知识的主要表达形式,也是研究数据的重要来源。然而,以叙述性文本形式表达的科学知识本身并不具备机器可重用性。为促进知识复用(例如用于综述研究),必须从文章中提取科学知识并在发表后将其组织到数据库中。由于手动完成这些活动耗时高且准确性不足,推动了知识自动提取技术的发展。本文以不同思路应对该问题,提出一种名为"重生"的预发表方法,确保科学知识在产生之初即可重用,即在知识生产过程中以机器可重用格式生成。我们采用开放研究知识图谱基础设施实现该方法,以实现FAIR原则下的科学知识组织。通过三个用例验证该方法,并探讨出版商与编辑在推广该方法中的作用。结果表明,与传统的手动及半自动化发表后提取技术相比,所提方法在知识丰富度、准确性及技术简洁性方面均具有显著优势。