The time and expense required to collect and label audio data has been a prohibitive factor in the availability of domain specific audio datasets. As the predictive specificity of a classifier depends on the specificity of the labels it is trained on, it follows that finely-labelled datasets are crucial for advances in machine learning. Aiming to stimulate progress in the field of machine listening, this paper introduces AeroSonicDB (YPAD-0523), a dataset of low-flying aircraft sounds for training acoustic detection and classification systems. This paper describes the method of exploiting ADS-B radio transmissions to passively collect and label audio samples. Provides a summary of the collated dataset. Presents baseline results from three binary classification models, then discusses the limitations of the current dataset and its future potential. The dataset contains 625 aircraft recordings ranging in event duration from 18 to 60 seconds, for a total of 8.87 hours of aircraft audio. These 625 samples feature 301 unique aircraft, each of which are supplied with 14 supplementary (non-acoustic) labels to describe the aircraft. The dataset also contains 3.52 hours of ambient background audio ("silence"), as a means to distinguish aircraft noise from other local environmental noises. Additionally, 6 hours of urban soundscape recordings (with aircraft annotations) are included as an ancillary method for evaluating model performance, and to provide a testing ground for real-time applications.
翻译:音频数据的采集与标注所需的时间和成本,已成为特定领域音频数据集可用性的制约因素。由于分类器的预测特异性取决于其训练标签的特异性,因此精细标注的数据集对于机器学习的进步至关重要。为促进机器听觉领域的发展,本文介绍了AeroSonicDB(YPAD-0523)数据集,该数据集包含低空飞行器声音样本,可用于训练声学检测与分类系统。本文阐述了利用ADS-B无线电传输被动采集并标注音频样本的方法,汇总了整理后的数据集,展示了三个二元分类模型的基线结果,并讨论了当前数据集的局限性及未来潜力。该数据集包含625条飞行器录音,事件时长从18秒到60秒不等,共计8.87小时的飞行器音频。这625个样本涵盖了301架不同的飞行器,每架飞行器均附有14个辅助(非声学)标签以描述其特征。此外,数据集还包含3.52小时的环境背景音频("静音"),用于区分飞行器噪声与其他本地环境噪声。同时,为评估模型性能及测试实时应用场景,另提供了6小时的城市声景录音(含飞行器标注)作为辅助方法。