YourTTS brings the power of a multilingual approach to the task of zero-shot multi-speaker TTS. Our method builds upon the VITS model and adds several novel modifications for zero-shot multi-speaker and multilingual training. We achieved state-of-the-art (SOTA) results in zero-shot multi-speaker TTS and results comparable to SOTA in zero-shot voice conversion on the VCTK dataset. Additionally, our approach achieves promising results in a target language with a single-speaker dataset, opening possibilities for zero-shot multi-speaker TTS and zero-shot voice conversion systems in low-resource languages. Finally, it is possible to fine-tune the YourTTS model with less than 1 minute of speech and achieve state-of-the-art results in voice similarity and with reasonable quality. This is important to allow synthesis for speakers with a very different voice or recording characteristics from those seen during training.
翻译:YourTTS将多语言方法的优势引入零样本多说话人TTS任务。我们的方法基于VITS模型,并增加了多项针对零样本多说话人与多语言训练的新改进。我们在VCTK数据集上的零样本多说话人TTS任务中取得了最先进(SOTA)的结果,同时在零样本语音转换任务中达到了与SOTA相当的性能。此外,我们的方法在单说话人数据集的目标语言中取得了有前景的结果,为低资源语言的零样本多说话人TTS和零样本语音转换系统开辟了可能性。最后,YourTTS模型可通过少于1分钟的语言数据微调,在语音相似度上获得最优结果,且质量合理。这对于合成具有与训练时显著不同的嗓音或录音特征的说话人语音至关重要。