We explore on various attention methods on frequency and channel dimensions for sound event detection (SED) in order to enhance performance with minimal increase in computational cost while leveraging domain knowledge to address the frequency dimension of audio data. We have introduced frequency dynamic convolution in a previous work to release the translational equivariance issue associated with 2D convolution on the frequency dimension of 2D audio data. Although this approach demonstrated state-of-the-art SED performance, it resulted in 2.5 times heavier model in terms of the number of parameters. To achieve comparable SED performance with computationally efficient methods to enhance practicality, we explore on lighter alternative attention methods. In addition, we focus of attention methods on frequency and channel dimensions as those are shown to be critical in SED. Joint application of SE modules on both frequency and channel dimension shows comparable performance to frequency dynamic convolution with only 2.7% increase in the model size compared to the baseline model. In addition, we performed class-wise comparison of various attention methods to further discuss their characteristics.
翻译:我们探究了多种在频率和通道维度上的注意力方法,用于声音事件检测(SED),旨在以最小的计算成本提升性能,同时利用领域知识处理音频数据的频率维度。在先前工作中,我们引入了频率动态卷积,以解决二维音频数据中频率维度上二维卷积存在的平移等变性不足问题。尽管该方法展现了最先进的SED性能,但它导致模型参数数量增加了2.5倍。为了在保持可比较的SED性能的同时,采用计算高效的方法以增强实用性,我们探索了更轻量的替代注意力方法。此外,我们重点关注频率和通道维度上的注意力方法,因为这些维度已被证明对SED至关重要。在频率和通道维度上联合应用SE模块,其性能可与频率动态卷积相媲美,且模型大小仅比基线模型增加2.7%。此外,我们还对多种注意力方法进行了类别级比较,以进一步讨论它们的特性。