Deep learning has been recently introduced for efficient acoustic howling suppression (AHS). However, the recurrent nature of howling creates a mismatch between offline training and streaming inference, limiting the quality of enhanced speech. To address this limitation, we propose a hybrid method that combines a Kalman filter with a self-attentive recurrent neural network (SARNN) to leverage their respective advantages for robust AHS. During offline training, a pre-processed signal obtained from the Kalman filter and an ideal microphone signal generated via teacher-forced training strategy are used to train the deep neural network (DNN). During streaming inference, the DNN's parameters are fixed while its output serves as a reference signal for updating the Kalman filter. Evaluation in both offline and streaming inference scenarios using simulated and real-recorded data shows that the proposed method efficiently suppresses howling and consistently outperforms baselines.
翻译:近年来,深度学习被引入用于高效声学啸叫抑制(AHS)。然而,啸叫的递归特性导致离线训练与流式推理之间存在失配,限制了增强语音的质量。为克服这一局限,我们提出一种将卡尔曼滤波与自注意力循环神经网络(SARNN)相结合的混合方法,以利用二者各自优势实现鲁棒的AHS。在离线训练阶段,使用卡尔曼滤波预处理信号与基于教师强制训练策略生成的理想麦克风信号共同训练深度神经网络(DNN)。在流式推理阶段,DNN参数固定不变,其输出作为参考信号用于更新卡尔曼滤波器。基于仿真与真实录音数据在离线与流式推理两种场景下的评估表明,所提方法能高效抑制啸叫,且性能持续优于现有基线方法。