Pre-training datasets, like ImageNet, have become the gold standard in medical image analysis. However, the emergence of self-supervised learning (SSL), which leverages unlabeled data to learn robust features, presents an opportunity to bypass the intensive labeling process. In this study, we explored if SSL for pre-training on non-medical images can be applied to chest radiographs and how it compares to supervised pre-training on non-medical images and on medical images. We utilized a vision transformer and initialized its weights based on (i) SSL pre-training on natural images (DINOv2), (ii) SL pre-training on natural images (ImageNet dataset), and (iii) SL pre-training on chest radiographs from the MIMIC-CXR database. We tested our approach on over 800,000 chest radiographs from six large global datasets, diagnosing more than 20 different imaging findings. Our SSL pre-training on curated images not only outperformed ImageNet-based pre-training (P<0.001 for all datasets) but, in certain cases, also exceeded SL on the MIMIC-CXR dataset. Our findings suggest that selecting the right pre-training strategy, especially with SSL, can be pivotal for improving artificial intelligence (AI)'s diagnostic accuracy in medical imaging. By demonstrating the promise of SSL in chest radiograph analysis, we underline a transformative shift towards more efficient and accurate AI models in medical imaging.
翻译:预训练数据集(如ImageNet)已成为医学图像分析领域的黄金标准。然而,基于无标注数据学习鲁棒特征的自监督学习(SSL)的出现,为绕过繁琐的标注过程提供了契机。本研究探究了使用非医学图像进行SSL预训练是否可应用于胸部X光片,并将其与基于非医学图像及医学图像的监督学习(SL)预训练方法进行了对比。我们采用视觉Transformer架构,分别基于以下策略初始化模型权重:(i)对自然图像进行SSL预训练(DINOv2),(ii)对自然图像进行SL预训练(ImageNet数据集),(iii)对来自MIMIC-CXR数据库的胸部X光片进行SL预训练。我们在来自六个大型全球数据集(超过80万张胸部X光片)上测试了该方法,共诊断20余种不同影像学表现。实验结果表明,基于精选图像的SSL预训练不仅优于基于ImageNet的预训练(所有数据集P<0.001),在某些情况下甚至超越了基于MIMIC-CXR数据集的SL方法。本研究表明,选择恰当的预训练策略(尤其是SSL方法)对于提升医疗影像人工智能(AI)诊断准确性具有关键作用。通过验证SSL在胸部X光片分析中的潜力,我们揭示了医疗影像领域向更高效、更精准AI模型转变的变革性趋势。