[Citation needed] Data usage and citation practices in medical imaging conferences

Medical imaging papers often focus on methodology, but the quality of the algorithms and the validity of the conclusions are highly dependent on the datasets used. As creating datasets requires a lot of effort, researchers often use publicly available datasets, there is however no adopted standard for citing the datasets used in scientific papers, leading to difficulty in tracking dataset usage. In this work, we present two open-source tools we created that could help with the detection of dataset usage, a pipeline \url{https://github.com/TheoSourget/Public_Medical_Datasets_References} using OpenAlex and full-text analysis, and a PDF annotation software \url{https://github.com/TheoSourget/pdf_annotator} used in our study to manually label the presence of datasets. We applied both tools on a study of the usage of 20 publicly available medical datasets in papers from MICCAI and MIDL. We compute the proportion and the evolution between 2013 and 2023 of 3 types of presence in a paper: cited, mentioned in the full text, cited and mentioned. Our findings demonstrate the concentration of the usage of a limited set of datasets. We also highlight different citing practices, making the automation of tracking difficult.

翻译：医学影像论文通常侧重于方法学，但算法质量和结论有效性高度依赖于所使用的数据集。由于创建数据集需要大量投入，研究者常使用公开数据集，然而目前尚无统一标准来规范科学论文中数据集的引用方式，导致难以追踪数据集的使用情况。本研究提出两款自主开发的开源工具：基于OpenAlex与全文分析的流水线工具（\url{https://github.com/TheoSourget/Public_Medical_Datasets_References}），以及用于人工标注数据集存在的PDF标注软件（\url{https://github.com/TheoSourget/pdf_annotator}）。我们应用这两款工具对MICCAI与MIDL会议论文中20个公开医学数据集的使用情况进行了分析，计算了2013至2023年间三类呈现方式（被引用、全文提及、既被引用又被提及）的比例及其演变趋势。研究结果表明，数据集使用呈现高度集中化特征，同时发现不同的引用实践对自动化追踪构成了显著挑战。

相关内容

数据集

关注 88

数据集，又称为资料集、数据集合或资料集合，是一种由数据所组成的集合。
Data set（或dataset）是一个数据的集合，通常以表格形式出现。每一列代表一个特定变量。每一行都对应于某一成员的数据集的问题。它列出的价值观为每一个变量，如身高和体重的一个物体或价值的随机数。每个数值被称为数据资料。对应于行数，该数据集的数据可能包括一个或多个成员。

《生成式模型: 变分自编码器与扩散模型》，75页ppt，Google DeepMind科学家Ruiqi Gao

专知会员服务

66+阅读 · 2023年6月10日

【CVPR 2022】一个完全无监督的框架，从噪声和部分测量中学习图像，Robust Equivariant Imaging: a fully unsupervised framework for learning to image

专知会员服务

25+阅读 · 2022年3月3日

生成性对抗网络:理论模型、评估指标和最近发展的概述，Generative Adversarial Networks (GANs): An Overview of Theoretical Model, Evaluation Metrics, and Recent Developments

专知会员服务

42+阅读 · 2020年5月30日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日