Automatic assessment of dysarthric speech is essential for sustained treatments and rehabilitation. However, obtaining atypical speech is challenging, often leading to data scarcity issues. To tackle the problem, we propose a novel automatic severity assessment method for dysarthric speech, using the self-supervised model in conjunction with multi-task learning. Wav2vec 2.0 XLS-R is jointly trained for two different tasks: severity classification and auxiliary automatic speech recognition (ASR). For the baseline experiments, we employ hand-crafted acoustic features and machine learning classifiers such as SVM, MLP, and XGBoost. Explored on the Korean dysarthric speech QoLT database, our model outperforms the traditional baseline methods, with a relative percentage increase of 1.25% for F1-score. In addition, the proposed model surpasses the model trained without ASR head, achieving 10.61% relative percentage improvements. Furthermore, we present how multi-task learning affects the severity classification performance by analyzing the latent representations and regularization effect.
翻译:针对构音障碍语音的自动评估是持续治疗与康复的关键环节。然而,获取非典型语音样本极具挑战性,常导致数据稀缺问题。为解决这一难题,我们提出了一种基于自监督模型结合多任务学习的构音障碍语音严重程度自动评估新方法。该方法联合训练Wav2vec 2.0 XLS-R模型完成两项任务:严重程度分类与辅助自动语音识别(ASR)。基线实验中,我们采用手工声学特征及SVM、MLP、XGBoost等机器学习分类器。基于韩国构音障碍语音QoLT数据库的测试表明,本模型优于传统基线方法,F1分数相对提升1.25%。此外,相较于未集成ASR分支的模型,所提模型实现了10.61%的相对性能提升。通过分析潜在表征与正则化效应,我们还揭示了多任务学习对严重程度分类性能的影响机制。