Unpaired text and audio injection have emerged as dominant methods for improving ASR performance in the absence of a large labeled corpus. However, little guidance exists on deploying these methods to improve production ASR systems that are trained on very large supervised corpora and with realistic requirements like a constrained model size and CPU budget, streaming capability, and a rich lattice for rescoring and for downstream NLU tasks. In this work, we compare three state-of-the-art semi-supervised methods encompassing both unpaired text and audio as well as several of their combinations in a controlled setting using joint training. We find that in our setting these methods offer many improvements beyond raw WER, including substantial gains in tail-word WER, decoder computation during inference, and lattice density.
翻译:无配对文本与音频注入已成为在缺乏大规模标注语料时提升语音识别性能的主流方法。然而,针对如何在受限于模型规模、CPU预算、流式处理能力以及适用于重评分与下游自然语言理解任务的丰富格结构的现实条件下,部署这些方法以改进基于极大规模监督语料训练的工业级语音识别系统,目前尚缺乏系统性指导。本研究在联合训练框架下,对三种涵盖无配对文本与音频的最先进半监督方法及其多种组合进行受控对比分析。结果表明,在该设定下,这些方法不仅能够降低原始词错误率,更能显著改善尾部词错误率、推理阶段解码器计算效率以及格结构密度。