Hybrid meetings have become increasingly necessary during the post-COVID period and also brought new challenges for solving audio-related problems. In particular, the interplay between acoustic echo and acoustic howling in a hybrid meeting makes the joint suppression of them difficult. This paper proposes a deep learning approach to tackle this problem by formulating a recurrent feedback suppression process as an instantaneous speech separation task using the teacher-forced training strategy. Specifically, a self-attentive recurrent neural network is utilized to extract the target speech from microphone recordings with accessible and learned reference signals, thus suppressing acoustic echo and acoustic howling simultaneously. Different combinations of input signals and loss functions have been investigated for performance improvement. Experimental results demonstrate the effectiveness of the proposed method for suppressing echo and howling jointly in hybrid meetings.
翻译:混合会议在后疫情时代变得日益必要,同时也为解决音频相关问题带来了新挑战。特别是,混合会议中声学回声与声学啸叫的相互作用使得对它们的联合抑制变得困难。本文提出了一种深度学习方法,通过采用教师强制训练策略,将循环反馈抑制过程表述为即时语音分离任务,从而解决该问题。具体而言,利用自注意力循环神经网络从可获取和已学习的参考信号的麦克风录音中提取目标语音,从而同时抑制声学回声和声学啸叫。研究了输入信号和损失函数的不同组合以提升性能。实验结果表明,所提方法在混合会议中联合抑制回声和啸叫的有效性。