Our previous work, the unified source-filter GAN (uSFGAN) vocoder, introduced a novel architecture based on the source-filter theory into the parallel waveform generative adversarial network to achieve high voice quality and pitch controllability. However, the high temporal resolution inputs result in high computation costs. Although the HiFi-GAN vocoder achieves fast high-fidelity voice generation thanks to the efficient upsampling-based generator architecture, the pitch controllability is severely limited. To realize a fast and pitch-controllable high-fidelity neural vocoder, we introduce the source-filter theory into HiFi-GAN by hierarchically conditioning the resonance filtering network on a well-estimated source excitation information. According to the experimental results, our proposed method outperforms HiFi-GAN and uSFGAN on a singing voice generation in voice quality and synthesis speed on a single CPU. Furthermore, unlike the uSFGAN vocoder, the proposed method can be easily adopted/integrated in real-time applications and end-to-end systems.
翻译:我们先前的工作,统一源-滤波器生成对抗网络(uSFGAN)声码器,将基于源-滤波器理论的新型架构引入并行波形生成对抗网络,以实现高语音质量和音高可控性。然而,高时间分辨率输入导致计算成本高昂。尽管HiFi-GAN声码器凭借高效的上采样生成器架构实现了快速高保真语音生成,但其音高可控性严重受限。为实现快速且音高可控的高保真神经声码器,我们通过将共振滤波网络基于良好估计的源激励信息进行层级条件化处理,将源-滤波器理论引入HiFi-GAN。实验结果表明,在单个CPU上,我们提出的方法在歌唱语音生成的语音质量和合成速度方面均优于HiFi-GAN和uSFGAN。此外,与uSFGAN声码器不同,所提方法可轻松应用于/集成到实时应用和端到端系统中。