In recent years, funding agencies and journals increasingly advocate for open science practices (e.g. data and method sharing) to improve the transparency, access, and reproducibility of science. However, quantifying these practices at scale has proven difficult. In this work, we leverage a large-scale dataset of 1.1M papers from arXiv that are representative of the fields of physics, math, and computer science to analyze the adoption of data and method link-sharing practices over time and their impact on article reception. To identify links to data and methods, we train a neural text classification model to automatically classify URL types based on contextual mentions in papers. We find evidence that the practice of link-sharing to methods and data is spreading as more papers include such URLs over time. Reproducibility efforts may also be spreading because the same links are being increasingly reused across papers (especially in computer science); and these links are increasingly concentrated within fewer web domains (e.g. Github) over time. Lastly, articles that share data and method links receive increased recognition in terms of citation count, with a stronger effect when the shared links are active (rather than defunct). Together, these findings demonstrate the increased spread and perceived value of data and method sharing practices in open science.
翻译:近年来,资助机构和期刊越来越多地倡导开放科学实践(例如数据与方法共享),以提高科学的透明度、可获取性和可重复性。然而,大规模量化这些实践一直具有挑战性。本研究利用来自arXiv的110万篇论文的大规模数据集(涵盖物理学、数学和计算机科学领域),分析数据与方法链接分享实践随时间的变化及其对文章接受度的影响。为识别数据和方法链接,我们训练了一个神经文本分类模型,基于论文中的上下文提及自动对URL类型进行分类。我们发现证据表明,方法和数据链接分享实践正在扩展:随时间推移,包含此类URL的论文数量增加。可重复性努力也可能在扩展,因为相同链接在论文间的重复使用率逐渐上升(尤其在计算机科学领域);并且这些链接随时间越来越集中于更少的网络域(如Github)。最后,分享数据与方法链接的文章在引用量上获得更高认可,且当共享链接处于活跃状态(而非失效)时,这一效应更强。综合而言,这些发现表明开放科学中数据与方法分享实践的普及及其感知价值的提升。