Acoustic echo cancellation (AEC), noise suppression (NS) and dereverberation (DR) are an integral part of modern full-duplex communication systems. As the demand for teleconferencing systems increases, addressing these tasks is required for an effective and efficient online meeting experience. Most prior research proposes solutions for these tasks separately, combining them with digital signal processing (DSP) based components, resulting in complex pipelines that are often impractical to deploy in real-world applications. This paper proposes a real-time cross-attention deep model, named DeepVQE, based on residual convolutional neural networks (CNNs) and recurrent neural networks (RNNs) to simultaneously address AEC, NS, and DR. We conduct several ablation studies to analyze the contributions of different components of our model to the overall performance. DeepVQE achieves state-of-the-art performance on non-personalized tracks from the ICASSP 2023 Acoustic Echo Cancellation Challenge and ICASSP 2023 Deep Noise Suppression Challenge test sets, showing that a single model can handle multiple tasks with excellent performance. Moreover, the model runs in real-time and has been successfully tested for the Microsoft Teams platform.
翻译:声学回声消除、噪声抑制与去混响是现代全双工通信系统的重要组成部分。随着远程会议系统需求的增长,高效且有效的在线会议体验需要同时解决这些任务。以往的研究大多针对这些任务分别提出解决方案,并结合基于数字信号处理的模块,导致复杂的流水线架构,在实际应用中往往难以部署。本文提出一种基于残差卷积神经网络与循环神经网络的实时交叉注意力深度模型DeepVQE,用于同时处理声学回声消除、噪声抑制与去混响任务。我们通过多项消融实验分析了模型各组件对整体性能的贡献。DeepVQE在ICASSP 2023声学回声消除挑战赛与ICASSP 2023深度噪声抑制挑战赛测试集的非个性化音轨上取得了当前最优性能,表明单一模型能够以卓越性能处理多项任务。此外,该模型可实现实时运行,并已成功通过Microsoft Teams平台的测试验证。