With fewer feature dimensions, filter banks are often used in light-weight full-band speech enhancement models. In order to further enhance the coarse speech in the sub-band domain, it is necessary to apply a post-filtering for harmonic retrieval. The signal processing-based comb filters used in RNNoise and PercepNet have limited performance and may cause speech quality degradation due to inaccurate fundamental frequency estimation. To tackle this problem, we propose a learnable comb filter to enhance harmonics. Based on the sub-band model, we design a DNN-based fundamental frequency estimator to estimate the discrete fundamental frequencies and a comb filter for harmonic enhancement, which are trained via an end-to-end pattern. The experiments show the advantages of our proposed method over PecepNet and DeepFilterNet.
翻译:由于特征维度较低,滤波器组常被用于轻量级全频带语音增强模型中。为了进一步增强子带域中的粗粒度语音,需采用后置滤波方法实现谐波恢复。RNNoise和PercepNet中基于信号处理的梳状滤波器因基频估计不准确,其性能存在局限且可能导致语音质量下降。针对该问题,我们提出一种可学习梳状滤波器以增强谐波。基于子带模型,我们设计了基于深度神经网络(DNN)的基频估计器用于离散基频估计,并构建了谐波增强梳状滤波器,两者通过端到端模式联合训练。实验结果表明,所提方法相较于PecepNet和DeepFilterNet具有显著优势。