Histopathology plays a central role in clinical medicine and biomedical research. While artificial intelligence shows promising results on many pathological tasks, generalization and dealing with rare diseases, where training data is scarce, remains a challenge. Distilling knowledge from unlabeled data into a foundation model before learning from, potentially limited, labeled data provides a viable path to address these challenges. In this work, we extend the state of the art of foundation models for digital pathology whole slide images by semi-automated data curation and incorporating pathologist domain knowledge. Specifically, we combine computational and pathologist domain knowledge (1) to curate a diverse dataset of 103k slides corresponding to 750 million image patches covering data from different fixation, staining, and scanning protocols as well as data from different indications and labs across the EU and US, (2) for grouping semantically similar slides and tissue patches, and (3) to augment the input images during training. We evaluate the resulting model on a set of public and internal benchmarks and show that although our foundation model is trained with an order of magnitude less slides, it performs on par or better than competing models. We expect that scaling our approach to more data and larger models will further increase its performance and capacity to deal with increasingly complex real world tasks in diagnostics and biomedical research.
翻译:组织病理学在临床医学和生物医学研究中占据核心地位。尽管人工智能在诸多病理学任务中展现出可喜成果,但模型泛化能力及训练数据稀缺的罕见疾病处理仍面临挑战。在利用潜在有限的标注数据进行学习之前,先从未标注数据中提取知识构建基础模型,为应对上述挑战提供了可行路径。本研究通过半自动化数据筛选并融入病理学家领域知识,拓展了数字病理学全切片图像基础模型的技术前沿。具体而言,我们结合计算与病理学领域知识实现了以下目标:(1) 筛选包含10.3万张切片(对应7.5亿个图像块)的多样化数据集,覆盖不同固定、染色和扫描方案处理的数据,以及来自欧盟和美国不同适应症及实验室的数据;(2) 对语义相似的切片和组织块进行分组;(3) 在训练过程中对输入图像进行增强。我们在公开及内部基准测试中评估了所提模型,结果表明:尽管该基础模型的训练切片数量较同类模型低一个数量级,但其性能仍可达到或超越竞品模型。我们预期,随着数据规模和模型参数量的扩大,本方法的性能及其应对诊断学与生物医学研究中日益复杂的现实任务的能力将进一步提升。