It has been shown that the intelligibility of noisy speech can be improved by speech enhancement (SE) algorithms. However, monaural SE has not been established as an effective frontend for automatic speech recognition (ASR) in noisy conditions compared to an ASR model trained on noisy speech directly. The divide between SE and ASR impedes the progress of robust ASR systems, especially as SE has made major advances in recent years. This paper focuses on eliminating this divide with an ARN (attentive recurrent network) time-domain and a CrossNet time-frequency domain enhancement models. The proposed systems fully decouple frontend enhancement and backend ASR trained only on clean speech. Results on the WSJ, CHiME-2, LibriSpeech, and CHiME-4 corpora demonstrate that ARN and CrossNet enhanced speech both translate to improved ASR results in noisy and reverberant environments, and generalize well to real acoustic scenarios. The proposed system outperforms the baselines trained on corrupted speech directly. Furthermore, it cuts the previous best word error rate (WER) on CHiME-2 by $28.4\%$ relatively with a $5.57\%$ WER, and achieves $3.32/4.44\%$ WER on single-channel CHiME-4 simulated/real test data without training on CHiME-4.
翻译:研究表明,语音增强(SE)算法可提升含噪语音的清晰度。然而,与直接基于含噪语音训练的自动语音识别(ASR)模型相比,单声道SE尚未被证明是噪声环境下ASR的有效前端。SE与ASR之间的鸿沟阻碍了鲁棒ASR系统的发展,尤其近年来SE领域已取得重大突破。本文聚焦于通过注意循环网络(ARN)时域增强模型与CrossNet时频域增强模型消除这一鸿沟。所提系统完全解耦前端增强与仅使用纯净语音训练的后端ASR。在WSJ、CHiME-2、LibriSpeech及CHiME-4语料库上的实验结果表明,ARN与CrossNet增强后的语音在噪声与混响环境中均能有效提升ASR性能,并具有良好的真实声学场景泛化能力。所提系统优于直接基于退化语音训练的基线模型,并在CHiME-2上以5.57%的词错误率(WER)相对降低先前最优WER达28.4%,同时在未使用CHiME-4训练数据的情况下,于单通道CHiME-4模拟/真实测试数据上分别取得3.32%/4.44%的WER。