Recent data- and learning-based sound source localization (SSL) methods have shown strong performance in challenging acoustic scenarios. However, little work has been done on adapting such methods to track consistently multiple sources appearing and disappearing, as would occur in reality. In this paper, we present a new training strategy for deep learning SSL models with a straightforward implementation based on the mean squared error of the optimal association between estimated and reference positions in the preceding time frames. It optimizes the desired properties of a tracking system: handling a time-varying number of sources and ordering localization estimates according to their trajectories, minimizing identity switches (IDSs). Evaluation on simulated data of multiple reverberant moving sources and on two model architectures proves its effectiveness on reducing identity switches without compromising frame-wise localization accuracy.
翻译:近年来,基于数据和学习的声源定位方法在具有挑战性的声学场景中表现出色。然而,针对这些方法在真实场景中适应多个声源出现和消失的持续跟踪问题的研究仍较少。本文提出一种新的深度学习声源定位模型训练策略,该策略基于前一时刻估计位置与参考位置最优关联的均方误差,实现方式简洁。该策略优化了跟踪系统的关键特性:处理时变数量的声源,并根据轨迹对定位估计进行排序,以最小化身份切换次数。基于多混响运动声源的仿真数据及两种模型架构的实验证明,该方法能在不降低帧级定位精度的前提下有效减少身份切换。