The digital transformation of the scientific publishing industry has led to dramatic improvements in content discoverability and information analytics. Unfortunately, these improvements have not been uniform across research areas. The scientific literature in the arts, humanities and social sciences (AHSS) still lags behind, in part due to the scale of analog backlogs, the persisting importance of national languages, and a publisher ecosystem made of many, small or medium enterprises. We propose a bottom-up approach to support publishers in creating and maintaining their own publication knowledge graphs in the open domain. We do so by releasing a pipeline able to extract structured information from the bibliographies and indexes of AHSS publications, disambiguate, normalize and export it as linked data. We test the proposed pipeline on Brill's Classics collection, and release an implementation in open source for further use and improvement.
翻译:科学出版业的数字化转型极大地提升了内容的可发现性与信息分析能力。然而,这些进步在不同研究领域间并不均衡。艺术、人文与社会科学领域的科学文献依然滞后,部分原因在于模拟时代文献积压的规模、国家语言的持续重要性,以及由众多中小型企业组成的出版生态系统。我们提出一种自下而上的方法,支持出版商在开放领域创建和维护自己的出版物知识图谱。为此,我们发布了一套流水线,能够从艺术、人文与社会科学出版物的参考书目与索引中提取结构化信息,进行消歧、规范化,并以关联数据形式导出。我们在布里尔经典著作集上测试了所提出的流水线,并以开源方式发布了实现代码,以供进一步使用与改进。