We propose a neural network model that can separate target speech sources from interfering sources at different angular regions using two microphones. The model is trained with simulated room impulse responses (RIRs) using omni-directional microphones without needing to collect real RIRs. By relying on specific angular regions and multiple room simulations, the model utilizes consistent time difference of arrival (TDOA) cues, or what we call delay contrast, to separate target and interference sources while remaining robust in various reverberation environments. We demonstrate the model is not only generalizable to a commercially available device with a slightly different microphone geometry, but also outperforms our previous work which uses one additional microphone on the same device. The model runs in real-time on-device and is suitable for low-latency streaming applications such as telephony and video conferencing.
翻译:我们提出一种神经网络模型,该模型利用两个麦克风即可从不同角度区域的干扰源中分离出目标语音源。该模型使用全向麦克风模拟的房间脉冲响应(RIR)进行训练,无需采集真实RIR。通过依赖特定角度区域和多种房间仿真,模型利用一致的到达时间差(TDOA)线索(即我们所谓的时延对比度)来分离目标源与干扰源,同时在各种混响环境中保持鲁棒性。我们证明该模型不仅能泛化至麦克风几何结构略有差异的商用设备,而且在相同设备上优于我们此前需额外使用一个麦克风的工作。该模型可在设备端实时运行,适用于电话会议、视频会议等低延迟流媒体应用场景。