Recently, deep learning-based beamforming algorithms have shown promising performance in target speech extraction tasks. However, most systems do not fully utilize spatial information. In this paper, we propose a target speech extraction network that utilizes spatial information to enhance the performance of neural beamformer. To achieve this, we first use the UNet-TCN structure to model input features and improve the estimation accuracy of the speech pre-separation module by avoiding information loss caused by direct dimensionality reduction in other models. Furthermore, we introduce a multi-head cross-attention mechanism that enhances the neural beamformer's perception of spatial information by making full use of the spatial information received by the array. Experimental results demonstrate that our approach, which incorporates a more reasonable target mask estimation network and a spatial information-based cross-attention mechanism into the neural beamformer, effectively improves speech separation performance.
翻译:近年来,基于深度学习的波束形成算法在目标语音提取任务中展现出良好性能。然而,大多数系统未能充分利用空间信息。本文提出一种利用空间信息增强神经波束形成器性能的目标语音提取网络。为此,我们首先采用UNet-TCN结构对输入特征进行建模,通过避免其他模型因直接降维造成的信息损失,提升语音预分离模块的估计精度。此外,我们引入多头交叉注意力机制,通过充分利用阵列接收的空间信息,增强神经波束形成器对空间信息的感知能力。实验结果表明,该方法将更合理的目标掩码估计网络和基于空间信息的交叉注意力机制融入神经波束形成器,有效提升了语音分离性能。