This paper presents a neural vocoder based on a denoising diffusion probabilistic model (DDPM) incorporating explicit periodic signals as auxiliary conditioning signals. Recently, DDPM-based neural vocoders have gained prominence as non-autoregressive models that can generate high-quality waveforms. The neural vocoders based on DDPM have the advantage of training with a simple time-domain loss. In practical applications, such as singing voice synthesis, there is a demand for neural vocoders to generate high-fidelity speech waveforms with flexible pitch control. However, conventional DDPM-based neural vocoders struggle to generate speech waveforms under such conditions. Our proposed model aims to accurately capture the periodic structure of speech waveforms by incorporating explicit periodic signals. Experimental results show that our model improves sound quality and provides better pitch control than conventional DDPM-based neural vocoders.
翻译:本文提出了一种基于去噪扩散概率模型(DDPM)的神经声码器,该模型将显式周期信号作为辅助条件信号。近年来,基于DDPM的神经声码器作为能够生成高质量波形的非自回归模型而备受关注。这类神经声码器具有使用简单时域损失进行训练的优势。在实际应用如歌声合成中,需要神经声码器能够在灵活的音高控制下生成高保真语音波形。然而,传统的基于DDPM的神经声码器难以在如此条件下生成语音波形。我们提出的模型通过引入显式周期信号,旨在精确捕捉语音波形的周期结构。实验结果表明,我们的模型在改善音质的同时,相比传统基于DDPM的神经声码器提供了更优的音高控制能力。