This study proposes a method based on fully convolutional neural networks (FCNs) to identify migratory birds from their songs, with the objective of recognizing which birds pass through certain areas and at what time. To determine the best FCN architecture, extensive experimentation was conducted through a grid search, exploring the optimal depth, width, and activation function of the network. The results showed that the optimal number of filters is 400 in the widest layer, with 4 convolutional blocks with maxpooling and an adaptive activation function. The proposed FCN offers a significant advantage over other techniques, as it can recognize the sound of a bird in audio of any length with an accuracy greater than 85%. Furthermore, due to its architecture, the network can detect more than one species from audio and can carry out near-real-time sound recognition. Additionally, the proposed method is lightweight, making it ideal for deployment and use in IoT devices. The study also presents a comparative analysis of the proposed method against other techniques, demonstrating an improvement of over 67% in the best-case scenario. These findings contribute to advancing the field of bird sound recognition and provide valuable insights into the practical application of FCNs in real-world scenarios.
翻译:本研究提出一种基于全卷积神经网络(FCN)的方法,通过鸟类鸣声识别候鸟,旨在确定特定区域中鸟类通过的种类及时间。为确定最佳FCN架构,通过网格搜索进行了大量实验,探索网络的最优深度、宽度及激活函数。结果表明,最宽层的最佳滤波器数量为400,包含4个带最大池化的卷积模块及自适应激活函数。该FCN相较其他技术具有显著优势,可识别任意时长音频中的鸟声,准确率超过85%。此外,由于其架构特点,网络能同时检测音频中的多个物种,并实现近实时声纹识别。所提方法轻量化,特别适合在物联网设备中部署应用。研究还将该方法与其它技术进行对比分析,显示在最佳场景下性能提升超过67%。这些成果推动了鸟声识别领域的发展,并为FCN在实际场景中的工程应用提供了重要参考。