Environmental sound recognition (ESR) is an emerging research topic in audio pattern recognition. Many tasks are presented to resort to computational models for ESR in real-life applications. However, current models are usually designed for individual tasks, and are not robust and applicable to other tasks. Cross-task models, which promote unified knowledge modeling across various tasks, have not been thoroughly investigated. In this article, we propose a cross-task model for three different tasks of ESR: 1) acoustic scene classification; 2) urban sound tagging; and 3) anomalous sound detection. An architecture named SE-Trans is presented that uses attention mechanism-based Squeeze-and-Excitation and Transformer encoder modules to learn the channelwise relationship and temporal dependencies of the acoustic features. FMix is employed as the data augmentation method that improves the performance of ESR. Evaluations for the three tasks are conducted on the recent databases of detection and classification of acoustic scenes and event challenges. The experimental results show that the proposed cross-task model achieves state-of-the-art performance on all tasks. Further analysis demonstrates that the proposed cross-task model can effectively utilize acoustic knowledge across different ESR tasks.
翻译:环境声音识别(ESR)是音频模式识别领域的新兴研究课题。在现实应用中,许多任务需借助计算模型实现ESR。然而,现有模型通常针对单一任务设计,缺乏鲁棒性且难以应用于其他任务。能够促进跨任务统一知识建模的跨任务模型尚未得到充分研究。本文针对ESR的三个不同任务提出跨任务模型:1)声学场景分类;2)城市声音标注;3)异常声音检测。提出名为SE-Trans的架构,该架构采用基于注意力机制的挤压激励模块与Transformer编码器模块,分别学习声学特征的通道关联与时序依赖关系。采用FMix作为数据增强方法提升ESR性能。基于声学场景与事件检测分类领域最新数据集对三项任务进行评测。实验结果表明,所提出的跨任务模型在所有任务上均达到最优性能。进一步分析表明,该跨任务模型能够有效利用不同ESR任务间的声学知识。