The goal of this paper is to generate realistic audio with a lightweight and fast diffusion-based vocoder, named FreGrad. Our framework consists of the following three key components: (1) We employ discrete wavelet transform that decomposes a complicated waveform into sub-band wavelets, which helps FreGrad to operate on a simple and concise feature space, (2) We design a frequency-aware dilated convolution that elevates frequency awareness, resulting in generating speech with accurate frequency information, and (3) We introduce a bag of tricks that boosts the generation quality of the proposed model. In our experiments, FreGrad achieves 3.7 times faster training time and 2.2 times faster inference speed compared to our baseline while reducing the model size by 0.6 times (only 1.78M parameters) without sacrificing the output quality. Audio samples are available at: https://mm.kaist.ac.kr/projects/FreGrad.
翻译:本文旨在通过一种轻量级且快速的基于扩散的声码器(命名为FreGrad)生成逼真音频。我们的框架包含以下三个关键组件:(1)采用离散小波变换将复杂波形分解为子带小波,有助于FreGrad在简洁的特征空间上运行;(2)设计了一种频率感知的膨胀卷积,提升频率敏感性,从而生成具有精确频率信息的语音;(3)引入了一套技巧集,以提高所提模型的生成质量。实验中,FreGrad的训练速度比基线快3.7倍,推理速度快2.2倍,同时模型大小缩减至0.6倍(仅含178万参数),且输出质量不受影响。音频样本见:https://mm.kaist.ac.kr/projects/FreGrad。