Audio tagging is an important task of mapping audio samples to their corresponding categories. Recently endeavours that exploit transformer models in this field have achieved great success. However, the quadratic self-attention cost limits the scaling of audio transformer models and further constrains the development of more universal audio models. In this paper, we attempt to solve this problem by proposing Audio Mamba, a self-attention-free approach that captures long audio spectrogram dependency with state space models. Our experimental results on two audio-tagging datasets demonstrate the parameter efficiency of Audio Mamba, it achieves comparable results to SOTA audio spectrogram transformers with one third parameters.
翻译:音频标记是将音频样本映射至对应类别的重要任务。近期在该领域应用Transformer模型的尝试已取得显著成功。然而,二次自注意力计算成本限制了音频Transformer模型的扩展性,并进一步制约了更通用音频模型的发展。本文试图通过提出Audio Mamba来解决该问题,这是一种基于状态空间模型的无自注意力方法,能够捕捉长音频频谱图的依赖关系。我们在两个音频标记数据集上的实验结果表明了Audio Mamba的参数效率——其仅需三分之一参数即可达到与最先进音频频谱图Transformer相当的性能。