This paper introduces an end-to-end neural speech restoration model, HD-DEMUCS, demonstrating efficacy across multiple distortion environments. Unlike conventional approaches that employ cascading frameworks to remove undesirable noise first and then restore missing signal components, our model performs these tasks in parallel using two heterogeneous decoder networks. Based on the U-Net style encoder-decoder framework, we attach an additional decoder so that each decoder network performs noise suppression or restoration separately. We carefully design each decoder architecture to operate appropriately depending on its objectives. Additionally, we improve performance by leveraging a learnable weighting factor, aggregating the two decoder output waveforms. Experimental results with objective metrics across various environments clearly demonstrate the effectiveness of our approach over a single decoder or multi-stage systems for general speech restoration task.
翻译:摘要:本文提出了一种端到端神经语音恢复模型HD-DEMUCS,在多种失真环境下均展现出卓越性能。不同于传统方法采用级联框架先去除噪声、再修复缺失信号成分的处理方式,本模型通过两个异构解码器网络并行执行上述任务。基于U-Net型编码器-解码器框架,我们新增一个解码器,使两个解码器网络分别独立执行噪声抑制或信号恢复任务。我们针对各解码器的目标精心设计了其网络架构。此外,通过引入可学习的加权因子对两个解码器的输出波形进行聚合,进一步提升了模型性能。在多环境下的客观指标实验结果表明,对于通用语音恢复任务,本方法相较单解码器系统或多阶段系统具有显著优势。