Accurately extracting clinical information from speech is critical to the diagnosis and treatment of many neurological conditions. As such, there is interest in leveraging AI for automatic, objective assessments of clinical speech to facilitate diagnosis and treatment of speech disorders. We explore transfer learning using foundation models, focusing on the impact of layer selection for the downstream task of predicting pathological speech features. We find that selecting an optimal layer can greatly improve performance (~15.8% increase in balanced accuracy per feature as compared to worst layer, ~13.6% increase as compared to final layer), though the best layer varies by predicted feature and does not always generalize well to unseen data. A learned weighted sum offers comparable performance to the average best layer in-distribution (only ~1.2% lower) and had strong generalization for out-of-distribution data (only 1.5% lower than the average best layer).
翻译:从语音中准确提取临床信息对于多种神经系统疾病的诊断与治疗至关重要。因此,利用人工智能对临床语音进行自动、客观的评估,以促进言语障碍的诊断和治疗,已成为研究热点。本文探索了基于基础模型的迁移学习方法,重点分析了层级选择对病理语音特征预测这一下游任务的影响。研究发现,选择最优层级可显著提升模型性能(与最差层级相比,每个特征的平衡准确率平均提升约15.8%;与最终层级相比平均提升约13.6%),但最佳层级因预测特征而异,且未必能良好泛化至未见数据。通过学习加权求和的方法,在分布内数据上取得了与平均最佳层级相当的性能(仅低约1.2%),并在分布外数据上表现出较强的泛化能力(仅比平均最佳层级低1.5%)。