Image segmentation holds a vital position in the realms of diagnosis and treatment within the medical domain. Traditional convolutional neural networks (CNNs) and Transformer models have made significant advancements in this realm, but they still encounter challenges because of limited receptive field or high computing complexity. Recently, State Space Models (SSMs), particularly Mamba and its variants, have demonstrated notable performance in the field of vision. However, their feature extraction methods may not be sufficiently effective and retain some redundant structures, leaving room for parameter reduction. Motivated by previous spatial and channel attention methods, we propose Triplet Mamba-UNet. The method leverages residual VSS Blocks to extract intensive contextual features, while Triplet SSM is employed to fuse features across spatial and channel dimensions. We conducted experiments on ISIC17, ISIC18, CVC-300, CVC-ClinicDB, Kvasir-SEG, CVC-ColonDB, and Kvasir-Instrument datasets, demonstrating the superior segmentation performance of our proposed TM-UNet. Additionally, compared to the previous VM-UNet, our model achieves a one-third reduction in parameters.
翻译:图像分割在医学领域的诊断与治疗中占据关键地位。传统卷积神经网络(CNN)和Transformer模型在该领域取得了显著进展,但仍因感受野受限或计算复杂度过高而面临挑战。近期,状态空间模型(SSM),特别是Mamba及其变体,在视觉领域展现出卓越性能。然而,其特征提取方法可能不够高效,且保留部分冗余结构,存在参数精简空间。受先前空间与通道注意力机制的启发,我们提出三元Mamba-UNet(Triplet Mamba-UNet)。该方法利用残差VSS模块提取密集上下文特征,同时通过三元SSM(Triplet SSM)融合空间与通道维度特征。我们在ISIC17、ISIC18、CVC-300、CVC-ClinicDB、Kvasir-SEG、CVC-ColonDB和Kvasir-Instrument数据集上开展实验,验证了所提TM-UNet的优越分割性能。此外,与先前VM-UNet相比,本模型参数量缩减三分之一。