Time-domain speech enhancement (SE) has recently been intensively investigated. Among recent works, DEMUCS introduces multi-resolution STFT loss to enhance performance. However, some resolutions used for STFT contain non-stationary signals, and it is challenging to learn multi-resolution frequency losses simultaneously with only one output. For better use of multi-resolution frequency information, we supplement multiple spectrograms in different frame lengths into the time-domain encoders. They extract stationary frequency information in both narrowband and wideband. We also adopt multiple decoder outputs, each of which computes its corresponding resolution frequency loss. Experimental results show that (1) it is more effective to fuse stationary frequency features than non-stationary features in the encoder, and (2) the multiple outputs consistent with the frequency loss improve performance. Experiments on the Voice-Bank dataset show that the proposed method obtained a 0.14 PESQ improvement.
翻译:时域语音增强(SE)近年来得到了广泛研究。在近期工作中,DEMUCS引入了多分辨率STFT损失以提升性能。然而,用于STFT的某些分辨率包含非平稳信号,且仅凭单一输出同时学习多分辨率频率损失颇具挑战。为更好地利用多分辨率频率信息,我们将不同帧长的多个频谱图补充到时域编码器中。这些编码器在窄带和宽带中均能提取平稳频率信息。我们还采用了多个解码器输出,每个输出计算其对应分辨率的频率损失。实验结果表明:(1) 在编码器中融合平稳频率特征比融合非平稳特征更为有效,(2) 与频率损失一致的多输出结构提升了性能。在Voice-Bank数据集上的实验显示,所提方法获得了0.14 PESQ的提升。