Neural network based approaches to speech enhancement have shown to be particularly powerful, being able to leverage a data-driven approach to result in a significant performance gain versus other approaches. Such approaches are reliant on artificially created labelled training data such that the neural model can be trained using intrusive loss functions which compare the output of the model with clean reference speech. Performance of such systems when enhancing real-world audio often suffers relative to their performance on simulated test data. In this work, a non-intrusive multi-metric prediction approach is introduced, wherein a model trained on artificial labelled data using inference of an adversarially trained metric prediction neural network. The proposed approach shows improved performance versus state-of-the-art systems on the recent CHiME-7 challenge \ac{UDASE} task evaluation sets.
翻译:基于神经网络的语音增强方法已被证明尤为强大,能够利用数据驱动的方法相比其他方法带来显著的性能提升。此类方法依赖于人工创建的带标签训练数据,以便通过将模型输出与干净参考语音进行比较的有损损失函数来训练神经模型。在增强真实音频时,此类系统的性能往往相较于其在模拟测试数据上的表现有所下降。本文提出了一种非侵入式的多度量预测方法,其中利用对抗训练度量预测神经网络的推断,对基于人工标签数据训练的模型进行优化。所提出的方法在近期CHiMe-7挑战赛\ac{UDASE}任务评估集上展现了优于现有最优系统的性能。