Domain adaptation using text-only corpus is challenging in end-to-end(E2E) speech recognition. Adaptation by synthesizing audio from text through TTS is resource-consuming. We present a method to learn Unified Speech-Text Representation in Conformer Transducer(USTR-CT) to enable fast domain adaptation using the text-only corpus. Different from the previous textogram method, an extra text encoder is introduced in our work to learn text representation and is removed during inference, so there is no modification for online deployment. To improve the efficiency of adaptation, single-step and multi-step adaptations are also explored. The experiments on adapting LibriSpeech to SPGISpeech show the proposed method reduces the word error rate(WER) by relatively 44% on the target domain, which is better than those of TTS method and textogram method. Also, it is shown the proposed method can be combined with internal language model estimation(ILME) to further improve the performance.
翻译:使用纯文本语料库进行领域自适应在端到端(E2E)语音识别中具有挑战性。通过文本转语音(TTS)从文本合成音频进行自适应需要耗费大量资源。我们提出了一种在Conformer换能器(USTR-CT)中学习统一语音-文本表示的方法,以实现利用纯文本语料库的快速领域自适应。与之前的textogram方法不同,我们的工作引入了一个额外的文本编码器来学习文本表示,并在推理阶段移除该编码器,因此无需修改在线部署。为提高自适应效率,我们还探索了单步和多步自适应。在将LibriSpeech自适应至SPGISpeech的实验表明,所提方法在目标领域上相对降低了44%的词错误率(WER),优于TTS方法和textogram方法。此外,所提方法还可与内部语言模型估计(ILME)结合以进一步提升性能。