We propose a new two-pass E2E speech recognition model that improves ASR performance by training on a combination of paired data and unpaired text data. Previously, the joint acoustic and text decoder (JATD) has shown promising results through the use of text data during model training and the recently introduced deliberation architecture has reduced recognition errors by leveraging first-pass decoding results. Our method, dubbed Deliberation-JATD, combines the spelling correcting abilities of deliberation with JATD's use of unpaired text data to further improve performance. The proposed model produces substantial gains across multiple test sets, especially those focused on rare words, where it reduces word error rate (WER) by between 12% and 22.5% relative. This is done without increasing model size or requiring multi-stage training, making Deliberation-JATD an efficient candidate for on-device applications.
翻译:我们提出了一种新的两遍端到端语音识别模型,通过结合配对数据和无配对文本数据的训练来提高ASR性能。此前,联合声学与文本解码器(JATD)通过在模型训练过程中使用文本数据展现了良好的效果,而最新引入的思辨架构则通过利用第一遍解码结果减少了识别错误。我们的方法命名为Deliberation-JATD,将思辨的拼写纠正能力与JATD利用无配对文本数据的特性相结合,以进一步提升性能。所提模型在多个测试集上,尤其是针对罕见词的测试集中取得了显著增益,将词错误率(WER)相对降低了12%至22.5%。这一改进无需增加模型大小或进行多阶段训练,使得Deliberation-JATD成为适用于设备端应用的高效候选方案。