Collecting high-quality studio recordings of audio is challenging, which limits the language coverage of text-to-speech (TTS) systems. This paper proposes a framework for scaling a multilingual TTS model to 100+ languages using found data without supervision. The proposed framework combines speech-text encoder pretraining with unsupervised training using untranscribed speech and unspoken text data sources, thereby leveraging massively multilingual joint speech and text representation learning. Without any transcribed speech in a new language, this TTS model can generate intelligible speech in >30 unseen languages (CER difference of <10% to ground truth). With just 15 minutes of transcribed, found data, we can reduce the intelligibility difference to 1% or less from the ground-truth, and achieve naturalness scores that match the ground-truth in several languages.
翻译:收集高质量录音室音频资源具有挑战性,限制了文本转语音系统的语言覆盖范围。本文提出一种框架,利用无监督方法在未标注数据基础上将多语种TTS模型扩展至100种以上语言。该框架结合语音-文本编码器预训练与无监督训练(使用未转录语音和无语音文本数据源),从而充分利用大规模多语种联合语音与文本表示学习。对于新语言中无任何转录语音的情况,该TTS模型可在30种以上未见语言中生成可理解语音(CER与真实值差距<10%)。仅需15分钟转录的"发现式"数据,即可将可理解性差异降至1%以内,并在多种语言中实现与真实值相当的自然度评分。