While recent large-scale text-to-speech (TTS) models have achieved significant progress, they still fall short in speech quality, similarity, and prosody. Considering speech intricately encompasses various attributes (e.g., content, prosody, timbre, and acoustic details) that pose significant challenges for generation, a natural idea is to factorize speech into individual subspaces representing different attributes and generate them individually. Motivated by it, we propose NaturalSpeech 3, a TTS system with novel factorized diffusion models to generate natural speech in a zero-shot way. Specifically, 1) we design a neural codec with factorized vector quantization (FVQ) to disentangle speech waveform into subspaces of content, prosody, timbre, and acoustic details; 2) we propose a factorized diffusion model to generate attributes in each subspace following its corresponding prompt. With this factorization design, NaturalSpeech 3 can effectively and efficiently model the intricate speech with disentangled subspaces in a divide-and-conquer way. Experiments show that NaturalSpeech 3 outperforms the state-of-the-art TTS systems on quality, similarity, prosody, and intelligibility. Furthermore, we achieve better performance by scaling to 1B parameters and 200K hours of training data.
翻译:尽管近期大规模文本转语音(TTS)模型取得了显著进展,但在语音质量、相似度及韵律方面仍存在不足。考虑到语音本身包含内容、韵律、音色及声学细节等多重属性,这些属性对生成任务构成了重大挑战,一种自然的思路是将语音分解为表示不同属性的独立子空间并分别生成。受此启发,我们提出NaturalSpeech 3——一种采用新型因子化扩散模型的TTS系统,能够以零样本方式生成自然语音。具体而言:1)设计了一种基于因子化向量量化(FVQ)的神经编解码器,将语音波形解耦为内容、韵律、音色和声学细节四个子空间;2)提出因子化扩散模型,根据各子空间对应的提示生成相应属性。通过这种因子化设计,NaturalSpeech 3能以分治策略利用解耦子空间高效建模复杂语音。实验表明,NaturalSpeech 3在质量、相似度、韵律和可理解性方面均优于现有最优TTS系统。此外,通过将模型参数量扩展至10亿、训练数据扩展至20万小时,我们进一步取得了更优性能。