We developed dysarthric speech intelligibility classifiers on 551,176 disordered speech samples contributed by a diverse set of 468 speakers, with a range of self-reported speaking disorders and rated for their overall intelligibility on a five-point scale. We trained three models following different deep learning approaches and evaluated them on ~94K utterances from 100 speakers. We further found the models to generalize well (without further training) on the TORGO database (100% accuracy), UASpeech (0.93 correlation), ALS-TDI PMP (0.81 AUC) datasets as well as on a dataset of realistic unprompted speech we gathered (106 dysarthric and 76 control speakers,~2300 samples). To advance research in this domain, we share one of our models at https://tfhub.dev/google/euphonia_spice/classification/1.
翻译:我们基于468名不同说话者贡献的551,176例构音障碍语音样本,开发了构音障碍语音清晰度分类器。这些样本涵盖一系列自我报告的口语障碍类型,并采用五分制对整体清晰度进行评分。我们训练了三种采用不同深度学习方法的模型,并在来自100名说话者的约9.4万条语句上进行了评估。进一步发现,这些模型(无需额外训练)在TORGO数据库(准确率100%)、UASpeech数据集(相关系数0.93)、ALS-TDI PMP数据集(AUC值0.81)以及我们收集的真实非提示语音数据集(包含106名构音障碍说话者和76名对照说话者,共约2300个样本)上均展现出良好的泛化能力。为推动该领域研究,我们已在https://tfhub.dev/google/euphonia_spice/classification/1上共享其中一个模型。