In this paper, we introduce a novel semi-supervised learning framework tailored for medical image segmentation. Central to our approach is the innovative Multi-scale Text-aware ViT-CNN Fusion scheme. This scheme adeptly combines the strengths of both ViTs and CNNs, capitalizing on the unique advantages of both architectures as well as the complementary information in vision-language modalities. Further enriching our framework, we propose the Multi-Axis Consistency framework for generating robust pseudo labels, thereby enhancing the semi-supervised learning process. Our extensive experiments on several widely-used datasets unequivocally demonstrate the efficacy of our approach.
翻译:本文提出了一种面向医学图像分割的新型半监督学习框架。该框架的核心在于创新的多尺度文本感知ViT-CNN融合方案,该方案巧妙结合了视觉变换器(ViT)与卷积神经网络(CNN)各自架构优势,同时充分利用了视觉-语言模态间的互补信息。为强化框架性能,我们进一步提出多轴一致性框架以生成稳健的伪标签,从而提升半监督学习效果。在多个广泛使用的数据集上的大量实验明确验证了本方法的有效性。