Biomedical image classification requires capturing of bio-informatics based on specific feature distribution. In most of such applications, there are mainly challenges due to limited availability of samples for diseased cases and imbalanced nature of dataset. This article presents the novel framework of multi-head self-attention for vision transformer (ViT) which makes capable of capturing the specific image features for classification and analysis. The proposed method uses the concept of residual connection for accumulating the best attention output in each block of multi-head attention. The proposed framework has been evaluated on two small datasets: (i) blood cell classification dataset and (ii) brain tumor detection using brain MRI images. The results show the significant improvement over traditional ViT and other convolution based state-of-the-art classification models.
翻译:生物医学图像分类需要基于特定的特征分布捕获生物信息学特征。在大多数此类应用中,主要面临因病变样本数量有限以及数据集分布不平衡所带来的挑战。本文提出了一种新颖的多头自注意力视觉Transformer(ViT)框架,能够捕获特定图像特征以进行分类与分析。所提出的方法利用残差连接的概念,在每个多头注意力模块中累积最优注意力输出。该框架已在两个小型数据集上进行了评估:(i)血细胞分类数据集和(ii)基于脑部MRI图像的脑肿瘤检测数据集。实验结果表明,与传统ViT及其他基于卷积的先进分类模型相比,该方法具有显著性能提升。