Localizing a moving sound source in the real world involves determining its direction-of-arrival (DOA) and distance relative to a microphone. Advancements in DOA estimation have been facilitated by data-driven methods optimized with large open-source datasets with microphone array recordings in diverse environments. In contrast, estimating a sound source's distance remains understudied. Existing approaches assume recordings by non-coincident microphones to use methods that are susceptible to differences in room reverberation. We present a CRNN able to estimate the distance of moving sound sources across multiple datasets featuring diverse rooms, outperforming a recently-published approach. We also characterize our model's performance as a function of sound source distance and different training losses. This analysis reveals optimal training using a loss that weighs model errors as an inverse function of the sound source true distance. Our study is the first to demonstrate that sound source distance estimation can be performed across diverse acoustic conditions using deep learning.
翻译:在现实世界中定位移动声源涉及确定其相对于麦克风的到达方向(DOA)和距离。数据驱动方法通过使用包含不同环境中麦克风阵列录音的大型开源数据集进行优化,促进了DOA估计的进展。相比之下,声源距离估计的研究仍不充分。现有方法假设使用非重合麦克风进行录音,所采用的技术易受室内混响差异的影响。我们提出了一种CRNN,能够在多个包含不同房间的数据集上估计移动声源的距离,其性能优于近期发表的方法。我们还刻画了模型性能随声源距离和不同训练损失函数的变化规律。分析表明,最优训练策略是使用一种损失函数,该损失以声源真实距离的反函数形式对模型误差进行加权。本研究首次证明,利用深度学习可在多样声学条件下实现声源距离估计。