Recent advancements in diffusion-based acoustic models have revolutionized data-sufficient single-speaker Text-to-Speech (TTS) approaches, with Grad-TTS being a prime example. However, diffusion models suffer from drift in training and sampling distributions due to imperfect score-matching. The sampling drift problem leads to these approaches struggling in multi-speaker scenarios in practice. In this paper, we present Multi-GradSpeech, a multi-speaker diffusion-based acoustic models which introduces the Consistent Diffusion Model (CDM) as a generative modeling approach. We enforce the consistency property of CDM during the training process to alleviate the sampling drift problem in the inference stage, resulting in significant improvements in multi-speaker TTS performance. Our experimental results corroborate that our proposed approach can improve the performance of different speakers involved in multi-speaker TTS compared to Grad-TTS, even outperforming the fine-tuning approach. Audio samples are available at https://welkinyang.github.io/multi-gradspeech/
翻译:近年来,基于扩散的声学模型在数据充足的单说话人文本转语音(TTS)领域取得了革命性进展,Grad-TTS便是其中典型代表。然而,由于评分匹配不完美,扩散模型在训练与采样分布之间存在漂移问题。在实际场景中,这种采样漂移问题导致相关方法难以有效处理多说话人TTS任务。本文提出Multi-GradSpeech——一种引入一致扩散模型(CDM)作为生成建模方法的多说话人扩散声学模型。我们在训练过程中强制施加CDM的一致性约束,以缓解推理阶段的采样漂移问题,从而显著提升多说话人TTS性能。实验结果表明,与Grad-TTS相比,本方法能够改善多说话人TTS中不同说话人的合成效果,甚至优于基于微调的方法。音频样本见https://welkinyang.github.io/multi-gradspeech/