Obtaining large pre-trained models that can be fine-tuned to new tasks with limited annotated samples has remained an open challenge for medical imaging data. While pre-trained deep networks on ImageNet and vision-language foundation models trained on web-scale data are prevailing approaches, their effectiveness on medical tasks is limited due to the significant domain shift between natural and medical images. To bridge this gap, we introduce LVM-Med, the first family of deep networks trained on large-scale medical datasets. We have collected approximately 1.3 million medical images from 55 publicly available datasets, covering a large number of organs and modalities such as CT, MRI, X-ray, and Ultrasound. We benchmark several state-of-the-art self-supervised algorithms on this dataset and propose a novel self-supervised contrastive learning algorithm using a graph-matching formulation. The proposed approach makes three contributions: (i) it integrates prior pair-wise image similarity metrics based on local and global information; (ii) it captures the structural constraints of feature embeddings through a loss function constructed via a combinatorial graph-matching objective; and (iii) it can be trained efficiently end-to-end using modern gradient-estimation techniques for black-box solvers. We thoroughly evaluate the proposed LVM-Med on 15 downstream medical tasks ranging from segmentation and classification to object detection, and both for the in and out-of-distribution settings. LVM-Med empirically outperforms a number of state-of-the-art supervised, self-supervised, and foundation models. For challenging tasks such as Brain Tumor Classification or Diabetic Retinopathy Grading, LVM-Med improves previous vision-language models trained on 1 billion masks by 6-7% while using only a ResNet-50.
翻译:获取能够针对有限标注样本的新任务进行微调的大规模预训练模型,仍是医学影像数据的开放挑战。尽管基于ImageNet的预训练深度网络和利用网络规模数据训练的语言-视觉基础模型是主流方法,但由于自然图像与医学图像之间存在显著的领域偏移,其医学任务效果有限。为弥合这一鸿沟,我们提出LVM-Med,这是首个在大规模医学数据集上训练的深度网络系列。我们收集了来自55个公开数据集的约130万张医学图像,覆盖CT、MRI、X光、超声等多种器官和模态。基于该数据集,我们评估了多种最先进的自监督算法,并提出了一种基于图匹配公式的新型自监督对比学习算法。该方法包含三项贡献:(i)整合基于局部和全局信息的先验成对图像相似性度量;(ii)通过基于组合图匹配目标的损失函数捕捉特征嵌入的结构约束;(iii)利用现代黑箱求解器的梯度估计技术实现高效的端到端训练。我们在15项下游医学任务(涵盖分割、分类及目标检测)中全面评估了所提出的LVM-Med,并同时考察了分布内与分布外场景。实验表明,LVM-Med在性能上超越了多种最先进的监督学习、自监督学习和基础模型。在脑肿瘤分类或糖尿病视网膜病变分级等挑战性任务中,LVM-Med仅使用ResNet-50架构,便将此前基于10亿掩码训练的视觉-语言模型提升了6-7%。