Spoken languages show significant variation across mandarin and accent. Despite the high performance of mandarin automatic speech recognition (ASR), accent ASR is still a challenge task. In this paper, we introduce meta-learning techniques for fast accent domain expansion in mandarin speech recognition, which expands the field of accents without deteriorating the performance of mandarin ASR. Meta-learning or learn-to-learn can learn general relation in multi domains not only for over-fitting a specific domain. So we select meta-learning in the domain expansion task. This more essential learning will cause improved performance on accent domain extension tasks. We combine the methods of meta learning and freeze of model parameters, which makes the recognition performance more stable in different cases and the training faster about 20%. Our approach significantly outperforms other methods about 3% relatively in the accent domain expansion task. Compared to the baseline model, it improves relatively 37% under the condition that the mandarin test set remains unchanged. In addition, it also proved this method to be effective on a large amount of data with a relative performance improvement of 4% on the accent test set.
翻译:口语在不同方言与普通话间呈现显著差异。尽管普通话自动语音识别(ASR)已取得高性能,但口音ASR仍是一项具有挑战性的任务。本文引入元学习技术,用于普通话语音识别中口音域的快速扩展,该方法在保持普通话ASR性能不受影响的前提下扩展口音覆盖范围。元学习(即学习如何学习)能够学习多领域的通用关联关系,而非仅针对特定领域进行过拟合。因此我们选择元学习技术处理域扩展任务。这种更具本质性的学习方式能够提升口音域扩展任务的性能。我们结合元学习方法与模型参数冻结策略,使识别性能在不同场景下更稳定,同时训练速度提升约20%。我们的方法在口音域扩展任务中相对其他方法显著提升约3%。与基线模型相比,在普通话测试集保持不变的情况下,性能相对提升37%。此外,在大规模数据集上的实验证明,该方法使口音测试集性能相对提升4%。