Transformers have demonstrated remarkable performance in natural language processing and computer vision. However, existing vision Transformers struggle to learn from limited medical data and are unable to generalize on diverse medical image tasks. To tackle these challenges, we present MedFormer, a data-scalable Transformer designed for generalizable 3D medical image segmentation. Our approach incorporates three key elements: a desirable inductive bias, hierarchical modeling with linear-complexity attention, and multi-scale feature fusion that integrates spatial and semantic information globally. MedFormer can learn across tiny- to large-scale data without pre-training. Comprehensive experiments demonstrate MedFormer's potential as a versatile segmentation backbone, outperforming CNNs and vision Transformers on seven public datasets covering multiple modalities (e.g., CT and MRI) and various medical targets (e.g., healthy organs, diseased tissues, and tumors). We provide public access to our models and evaluation pipeline, offering solid baselines and unbiased comparisons to advance a wide range of downstream clinical applications.
翻译:Transformer在自然语言处理和计算机视觉中展现了卓越的性能。然而,现有视觉Transformer难以从有限医学数据中学习,且无法泛化至多样化的医学图像任务。为应对这些挑战,我们提出了MedFormer——一种面向可泛化三维医学图像分割的数据可扩展Transformer。该方法整合了三个关键要素:理想的归纳偏置、具有线性复杂度注意力机制的层级建模,以及融合全局空间与语义信息的多尺度特征融合。MedFormer无需预训练即可在极小到大规模数据上学习。综合实验表明,MedFormer具有作为通用分割骨干网络的潜力,在涵盖多种模态(如CT和MRI)及多种医学目标(如健康器官、病变组织和肿瘤)的七个公开数据集上,其性能优于CNN和视觉Transformer。我们已公开提供模型与评估流程,为推进广泛的临床下游应用提供稳健基线及无偏比较。