3D speech enhancement has attracted much attention in recent years with the development of augmented reality technology. Traditional denoising convolutional autoencoders have limitations in extracting dynamic voice information. In this paper, we propose a two-stage autoencoder neural network for 3D speech enhancement. We incorporate a dual-path recurrent neural network block into the convolutional autoencoder to iteratively apply time-domain and frequency-domain modeling in an alternate fashion. And an attention mechanism for fusing the high-dimension features is proposed. We also introduce a loss function to simultaneously optimize the network in the time-frequency and time domains. Experimental results show that our system outperforms the state-of-the-art systems on the dataset of ICASSP L3DAS23 challenge.
翻译:近年来,随着增强现实技术的发展,三维语音增强受到了广泛关注。传统的去噪卷积自编码器在提取动态语音信息方面存在局限性。本文提出了一种用于三维语音增强的两阶段自编码器神经网络。我们在卷积自编码器中引入双路径循环神经网络模块,以交替方式迭代地进行时域和频域建模,并提出了一个用于融合高维特征的注意力机制。同时,我们引入了一种损失函数,以同时在时频域和时域优化网络。实验结果表明,在ICASSP L3DAS23挑战赛数据集上,我们的系统性能优于现有最先进系统。