In recent years, self-supervised learning (SSL) models have produced promising results in a variety of speech-processing tasks, especially in contexts of data scarcity. In this paper, we study the use of SSL models for the task of mispronunciation detection for second language learners. We compare two downstream approaches: 1) training the model for phone recognition (PR) using native English data, and 2) training a model directly for the target task using non-native English data. We compare the performance of these two approaches for various SSL representations as well as a representation extracted from a traditional DNN-based speech recognition model. We evaluate the models on L2Arctic and EpaDB, two datasets of non-native speech annotated with pronunciation labels at the phone level. Overall, we find that using a downstream model trained for the target task gives the best performance and that most upstream models perform similarly for the task.
翻译:近年来,自监督学习模型在多种语音处理任务中展现出可喜的成果,尤其在数据稀缺场景下表现突出。本文研究了自监督模型在第二语言学习者错误发音检测任务中的应用。我们比较了两种下游方法:1)使用英语母语数据训练音素识别模型;2)使用非母语英语数据直接训练目标任务模型。我们对比了这两种方法在不同自监督表示以及传统深度神经网络语音识别模型提取的表示上的性能。评估采用L2Arctic和EpaDB两个标注了音素级别发音标签的非母语语音数据集。总体而言,我们发现针对目标任务训练的上下游模型性能最优,且大多数上游模型在该任务上表现相近。