In recent years, Transformer-based models such as the Switch Transformer have achieved remarkable results in natural language processing tasks. However, these models are often too complex and require extensive pre-training, which limits their effectiveness for small clinical text classification tasks with limited data. In this study, we propose a simplified Switch Transformer framework and train it from scratch on a small French clinical text classification dataset at CHU Sainte-Justine hospital. Our results demonstrate that the simplified small-scale Transformer models outperform pre-trained BERT-based models, including DistillBERT, CamemBERT, FlauBERT, and FrALBERT. Additionally, using a mixture of expert mechanisms from the Switch Transformer helps capture diverse patterns; hence, the proposed approach achieves better results than a conventional Transformer with the self-attention mechanism. Finally, our proposed framework achieves an accuracy of 87\%, precision at 87\%, and recall at 85\%, compared to the third-best pre-trained BERT-based model, FlauBERT, which achieved an accuracy of 84\%, precision at 84\%, and recall at 84\%. However, Switch Transformers have limitations, including a generalization gap and sharp minima. We compare it with a multi-layer perceptron neural network for small French clinical narratives classification and show that the latter outperforms all other models.
翻译:近年来,基于Transformer的模型(如Switch Transformer)在自然语言处理任务中取得了显著成果。然而,这类模型通常过于复杂且需要大规模预训练,这限制了其在数据量有限的少量临床文本分类任务中的有效性。在本研究中,我们提出了一种简化的Switch Transformer框架,并在CHU Sainte-Justine医院的小型法语临床文本分类数据集上从头开始训练。实验结果表明,简化的中小型Transformer模型优于基于预训练BERT的模型,包括DistillBERT、CamemBERT、FlauBERT和FrALBERT。此外,利用Switch Transformer中的专家混合机制有助于捕获多样化的模式;因此,所提出的方法比具有自注意力机制的传统Transformer取得了更好的结果。最终,所提出的框架达到了87%的准确率、87%的精确率和85%的召回率,而排名第三的预训练BERT模型FlauBERT取得了84%的准确率、84%的精确率和84%的召回率。然而,Switch Transformer存在局限性,包括泛化差距和尖锐最小值。我们将其与多层感知器神经网络在小型法语临床叙述分类任务中进行比较,发现后者优于所有其他模型。