In this work, we present DiffVoice, a novel text-to-speech model based on latent diffusion. We propose to first encode speech signals into a phoneme-rate latent representation with a variational autoencoder enhanced by adversarial training, and then jointly model the duration and the latent representation with a diffusion model. Subjective evaluations on LJSpeech and LibriTTS datasets demonstrate that our method beats the best publicly available systems in naturalness. By adopting recent generative inverse problem solving algorithms for diffusion models, DiffVoice achieves the state-of-the-art performance in text-based speech editing, and zero-shot adaptation.
翻译:本文提出DiffVoice,一种基于潜在扩散的新型文本转语音模型。我们首先通过对抗训练增强的变分自编码器将语音信号编码为音素率的潜在表示,然后利用扩散模型联合建模持续时间和潜在表示。在LJSpeech和LibriTTS数据集上的主观评估表明,我们的方法在自然度上超越了现有最佳公开系统。通过采用面向扩散模型的最新生成式逆问题求解算法,DiffVoice在基于文本的语音编辑和零样本自适应任务中达到了最先进的性能。