Passive acoustic monitoring enables large-scale observation of wildlife, but most bioacoustic classifiers only predict species presence in a time window without localizing vocalizations precisely in time or frequency, limiting downstream analyses. We formulate bird vocalization detection as an object detection task on spectrograms and train YOLO11 models to localize bird calls in dense tropical soundscapes from Singapore. We additionally introduce an open-source browser-based annotation tool and propose Intersection over Minimum (IoMin), an evaluation metric that better handles ambiguous acoustic boundaries than standard IoU and is better suited to the problem at hand. The best YOLO model nearly doubles baseline performance on in-distribution soundscapes from Singapore (81.8% vs. 42.1% IoMin@50 F1-score) while still outperforming the baseline on unseen out-of-distribution recordings from Hawaii (58.6% vs. 48.6%). These results suggest that object detection frameworks are a promising approach to time-frequency localization of animal vocalizations in complex soundscapes.
翻译:被动声学监测能够实现对野生动物的规模化观测,但大多数生物声学分类器仅能在时间窗口内预测物种存在,而无法在时间或频率上精确定位发声信号,限制了后续分析。本研究将鸟类发声检测形式化为声谱图上的目标检测任务,并训练YOLO11模型对新加坡密集热带声景中的鸟类叫声进行定位。我们额外引入了一款基于浏览器的开源标注工具,并提出了交叠最小值(Intersection over Minimum,IoMin)评估指标,该指标相比标准IoU能更好地处理模糊声学边界,更适用于当前问题。最佳YOLO模型在新加坡域内声景上的基线性能提升近一倍(81.8%对比42.1% IoMin@50 F1分数),同时在未见过的夏威夷域外录音上仍优于基线(58.6%对比48.6%)。这些结果表明,目标检测框架是解决复杂声景中动物发声时频定位问题的一种有前景的方法。