In this study, we present a method for generating automated anatomy segmentation datasets using a sequential process that involves nnU-Net-based pseudo-labeling and anatomy-guided pseudo-label refinement. By combining various fragmented knowledge bases, we generate a dataset of whole-body CT scans with $142$ voxel-level labels for 533 volumes providing comprehensive anatomical coverage which experts have approved. Our proposed procedure does not rely on manual annotation during the label aggregation stage. We examine its plausibility and usefulness using three complementary checks: Human expert evaluation which approved the dataset, a Deep Learning usefulness benchmark on the BTCV dataset in which we achieve 85% dice score without using its training dataset, and medical validity checks. This evaluation procedure combines scalable automated checks with labor-intensive high-quality expert checks. Besides the dataset, we release our trained unified anatomical segmentation model capable of predicting $142$ anatomical structures on CT data.
翻译:在本研究中,我们提出了一种通过顺序流程自动生成解剖分割数据集的方法,该流程包括基于nnU-Net的伪标注和解剖引导的伪标注细化。通过整合各种碎片化的知识库,我们生成了一个全身CT扫描数据集,其中包含533个扫描体积的142个体素级标签,提供了专家认可的全面解剖覆盖。所提出的流程在标签聚合阶段不依赖人工标注。我们通过三种互补的验证手段检验了其合理性和实用性:专家评估(认可了该数据集)、深度学习在BTCV数据集上的有用性基准测试(在不使用其训练数据集的情况下实现了85%的Dice分数)以及医学有效性检查。这一评估流程结合了可扩展的自动化检查与劳动密集型的高质量专家检查。除数据集外,我们还发布了训练好的统一解剖分割模型,该模型能够预测CT数据中的142个解剖结构。