Despite huge improvements in automatic speech recognition (ASR) employing neural networks, ASR systems still suffer from a lack of robustness and generalizability issues due to domain shifting. This is mainly because principal corpus design criteria are often not identified and examined adequately while compiling ASR datasets. In this study, we investigate the robustness of the state-of-the-art transfer learning approaches such as self-supervised wav2vec 2.0 and weakly supervised Whisper as well as fully supervised convolutional neural networks (CNNs) for multi-domain ASR. We also demonstrate the significance of domain selection while building a corpus by assessing these models on a novel multi-domain Bangladeshi Bangla ASR evaluation benchmark - BanSpeech, which contains approximately 6.52 hours of human-annotated speech and 8085 utterances from 13 distinct domains. SUBAK.KO, a mostly read speech corpus for the morphologically rich language Bangla, has been used to train the ASR systems. Experimental evaluation reveals that self-supervised cross-lingual pre-training is the best strategy compared to weak supervision and full supervision to tackle the multi-domain ASR task. Moreover, the ASR models trained on SUBAK.KO face difficulty recognizing speech from domains with mostly spontaneous speech. The BanSpeech will be publicly available to meet the need for a challenging evaluation benchmark for Bangla ASR.
翻译:尽管利用神经网络的自动语音识别(ASR)技术取得了巨大进步,但由于领域迁移问题,ASR系统仍面临鲁棒性和泛化能力不足的挑战。这主要是因为构建ASR数据集时,通常未能充分识别和检验核心语料设计标准。本研究探究了自监督wav2vec 2.0、弱监督Whisper以及全监督卷积神经网络(CNN)等先进迁移学习方法在多领域ASR中的鲁棒性。我们通过评估这些模型在新型多领域孟加拉国孟加拉语ASR评测基准——BanSpeech(包含约6.52小时人工标注语音和来自13个不同领域的8085条话语)上的表现,进一步证明了语料构建中领域选择的重要性。本文采用SUBAK.KO(一个主要为针对形态丰富的孟加拉语开发的朗读语音语料库)来训练ASR系统。实验评估表明,与弱监督和全监督方法相比,自监督跨语言预训练是应对多领域ASR任务的最佳策略。此外,基于SUBAK.KO训练的ASR模型在识别包含大量自发音语的领域语音时存在困难。为满足孟加拉语ASR对挑战性评测基准的需求,BanSpeech将公开发布。