Time-variant factors often occur in real-world full-duplex communication applications. Some of them are caused by the complex environment such as non-stationary environmental noises and varying acoustic path while some are caused by the communication system such as the dynamic delay between the far-end and near-end signals. Current end-to-end deep neural network (DNN) based methods usually model the time-variant components implicitly and can hardly handle the unpredictable time-variance in real-time speech enhancement. To explicitly capture the time-variant components, we propose a dynamic kernel generation (DKG) module that can be introduced as a learnable plug-in to a DNN-based end-to-end pipeline. Specifically, the DKG module generates a convolutional kernel regarding to each input audio frame, so that the DNN model is able to dynamically adjust its weights according to the input signal during inference. Experimental results verify that DKG module improves the performance of the model under time-variant scenarios, in the joint acoustic echo cancellation (AEC) and deep noise suppression (DNS) tasks.
翻译:时变因素常见于实际全双工通信应用中,部分由复杂环境(如非平稳环境噪声和时变声学路径)引起,部分由通信系统(如远端与近端信号间的动态延迟)导致。当前基于端到端深度神经网络的方法通常隐式建模时变分量,难以应对实时语音增强中不可预测的时变性。为显式捕捉时变分量,我们提出动态核生成模块,可作为可学习插件引入基于深度神经网络的端到端流水线。具体而言,动态核生成模块为每个输入音频帧生成卷积核,使深度神经网络模型在推理过程中能根据输入信号动态调整其权重。实验结果表明,在联合声学回声消除与深度噪声抑制任务中,动态核生成模块提升了模型在时变场景下的性能表现。