High quality speech capture has been widely studied for both voice communication and human computer interface reasons. To improve the capture performance, we can often find multi-microphone speech enhancement techniques deployed on various devices. Multi-microphone speech enhancement problem is often decomposed into two decoupled steps: a beamformer that provides spatial filtering and a single-channel speech enhancement model that cleans up the beamformer output. In this work, we propose a speech enhancement solution that takes both the raw microphone and beamformer outputs as the input for an ML model. We devise a simple yet effective training scheme that allows the model to learn from the cues of the beamformer by contrasting the two inputs and greatly boost its capability in spatial rejection, while conducting the general tasks of denoising and dereverberation. The proposed solution takes advantage of classical spatial filtering algorithms instead of competing with them. By design, the beamformer module then could be selected separately and does not require a large amount of data to be optimized for a given form factor, and the network model can be considered as a standalone module which is highly transferable independently from the microphone array. We name the ML module in our solution as GSENet, short for Guided Speech Enhancement Network. We demonstrate its effectiveness on real world data collected on multi-microphone devices in terms of the suppression of noise and interfering speech.
翻译:高质量的语音捕获技术因语音通信与人机交互的需求而得到广泛研究。为提升捕获性能,我们常能在各类设备上部署多麦克风语音增强技术。多麦克风语音增强问题通常分解为两个解耦步骤:提供空间滤波的波束成形器,以及用于净化波束成形器输出的单通道语音增强模型。本研究提出一种语音增强解决方案,将原始麦克风信号与波束成形器输出同时作为机器学习模型的输入。我们设计了一种简洁高效的训练方案,使模型能够通过对比两类输入信号学习波束成形器的空间线索,在完成去噪与去混响等常规任务的同时,显著提升其空间抑制能力。该方案并非与经典空间滤波算法竞争,而是利用其优势。通过这种设计,波束成形器模块可独立选择,无需针对特定形态参数进行大量数据优化;而网络模型可作为独立模块,具备跨麦克风阵列的高迁移性。我们将该方案中的机器学习模块命名为GSENet(Guided Speech Enhancement Network)。基于多麦克风设备采集的真实数据实验表明,该模型在抑制噪声与干扰语音方面具有显著效果。