There has been exploding interest in embracing Transformer-based architectures for medical image segmentation. However, the lack of large-scale annotated medical datasets make achieving performances equivalent to those in natural images challenging. Convolutional networks, in contrast, have higher inductive biases and consequently, are easily trainable to high performance. Recently, the ConvNeXt architecture attempted to modernize the standard ConvNet by mirroring Transformer blocks. In this work, we improve upon this to design a modernized and scalable convolutional architecture customized to challenges of data-scarce medical settings. We introduce MedNeXt, a Transformer-inspired large kernel segmentation network which introduces - 1) A fully ConvNeXt 3D Encoder-Decoder Network for medical image segmentation, 2) Residual ConvNeXt up and downsampling blocks to preserve semantic richness across scales, 3) A novel technique to iteratively increase kernel sizes by upsampling small kernel networks, to prevent performance saturation on limited medical data, 4) Compound scaling at multiple levels (depth, width, kernel size) of MedNeXt. This leads to state-of-the-art performance on 4 tasks on CT and MRI modalities and varying dataset sizes, representing a modernized deep architecture for medical image segmentation.
翻译:近年来,基于Transformer架构的医学图像分割方法引发了极大关注。然而,大规模标注医学数据集的匮乏使得实现与自然图像领域相当的性能面临挑战。相比之下,卷积网络具有更强的归纳偏置,因此更易训练至高性能。近期,ConvNeXt架构通过模仿Transformer模块尝试对标准卷积网络进行现代化改造。在本工作中,我们对此进行改进,设计了一种针对医学数据稀缺场景挑战的现代化、可扩展卷积架构。我们提出受Transformer启发的大核分割网络MedNeXt,其包含:1)用于医学图像分割的完整ConvNeXt 3D编码器-解码器网络;2)残差ConvNeXt上采样和下采样模块,以保持跨尺度的语义丰富性;3)一种通过上采样小核网络逐步增大卷积核大小的新技术,以规避有限医学数据上的性能饱和问题;4)MedNeXt在深度、宽度、核大小等多层次上的复合缩放。该方法在CT和MRI模态的四个任务及不同数据集规模上均达到最优性能,代表了医学图像分割的现代化深度架构。