Learning-based methods have become ubiquitous in speaker localization. Existing systems rely on simulated training sets for the lack of sufficiently large, diverse and annotated real datasets. Most room acoustics simulators used for this purpose rely on the image source method (ISM) because of its computational efficiency. This paper argues that carefully extending the ISM to incorporate more realistic surface, source and microphone responses into training sets can significantly boost the real-world performance of speaker localization systems. It is shown that increasing the training-set realism of a state-of-the-art direction-of-arrival estimator yields consistent improvements across three different real test sets featuring human speakers in a variety of rooms and various microphone arrays. An ablation study further reveals that every added layer of realism contributes positively to these improvements.
翻译:基于学习的方法在发言者定位中已变得无处不在。现有系统依赖模拟训练集,因为缺乏足够大、多样化且带有标注的真实数据集。为此目的使用的大多数室内声学模拟器依赖镜像源方法(ISM),因其计算效率高。本文论证,通过仔细扩展ISM以将更真实的表面、声源和麦克风响应纳入训练集,可以显著提升发言者定位系统在实际场景中的性能。研究表明,提升最先进的到达方向估计器的训练集真实度,能在三个包含不同房间和多种麦克风阵列的真实测试集上带来一致性的改进。一项消融研究进一步揭示,每一层新增的真实性都有助于这些改进。