We propose an end-to-end ASR system that can be trained on transcribed speech data, text data, or a mixture of both. For text-only training, our extended ASR model uses an integrated auxiliary TTS block that creates mel spectrograms from the text. This block contains a conventional non-autoregressive text-to-mel-spectrogram generator augmented with a GAN enhancer to improve the spectrogram quality. The proposed system can improve the accuracy of the ASR model on a new domain by using text-only data, and allows to significantly surpass conventional audio-text training by using large text corpora.
翻译:我们提出一种端到端自动语音识别(ASR)系统,该系统可基于转录语音数据、文本数据或两者的混合进行训练。针对纯文本训练,我们的扩展ASR模型采用集成的辅助文本转语音(TTS)模块,该模块可从文本生成梅尔频谱图。该模块包含一个经GAN增强器改进频谱图质量的常规非自回归文本到梅尔频谱图生成器。所提系统通过仅使用文本数据即可提升ASR模型在新领域上的精度,并利用大规模文本语料库显著超越传统的音频-文本联合训练方法。