We introduce a bilingual solution to support English as secondary locale for most primary locales in hybrid automatic speech recognition (ASR) settings. Our key developments constitute: (a) pronunciation lexicon with grapheme units instead of phone units, (b) a fully bilingual alignment model and subsequently bilingual streaming transformer model, (c) a parallel encoder structure with language identification (LID) loss, (d) parallel encoder with an auxiliary loss for monolingual projections. We conclude that in comparison to LID loss, our proposed auxiliary loss is superior in specializing the parallel encoders to respective monolingual locales, and that contributes to stronger bilingual learning. We evaluate our work on large-scale training and test tasks for bilingual Spanish (ES) and bilingual Italian (IT) applications. Our bilingual models demonstrate strong English code-mixing capability. In particular, the bilingual IT model improves the word error rate (WER) for a code-mix IT task from 46.5% to 13.8%, while also achieving a close parity (9.6%) with the monolingual IT model (9.5%) over IT tests.
翻译:我们提出了一种双语解决方案,用于在混合自动语音识别(ASR)设置中为大多数主语言地区支持英语作为第二语言。我们的关键进展包括:(a)采用字素单元而非音素单元的发音词典,(b)完全双语的声学对齐模型及随后的双语流式Transformer模型,(c)带有语言识别(LID)损失的并行编码器结构,(d)带有单语言投影辅助损失的并行编码器。我们得出结论:与LID损失相比,我们提出的辅助损失在使并行编码器专注于各自单语言区域方面表现更优,从而有助于更强的双语学习。我们在大规模训练和测试任务上评估了双语西班牙语(ES)和双语意大利语(IT)应用的效果。我们的双语模型展现了强大的英语代码混合能力。具体而言,双语IT模型将代码混合IT任务的词错误率(WER)从46.5%降低到13.8%,同时在IT测试上实现了与单语言IT模型(9.5%)接近的同等水平(9.6%)。