In this paper, we propose SemanticAC, a semantics-assisted framework for Audio Classification to better leverage the semantic information. Unlike conventional audio classification methods that treat class labels as discrete vectors, we employ a language model to extract abundant semantics from labels and optimize the semantic consistency between audio signals and their labels. We verify that simple textual information from labels and advanced pretraining models enable more abundant semantic supervision for better performance. Specifically, we design a text encoder to capture the semantic information from the text extension of labels. Then we map the audio signals to align with the semantics of corresponding class labels via an audio encoder and a similarity calculation module so as to enforce the semantic consistency. Extensive experiments on two audio datasets, ESC-50 and US8K demonstrate that our proposed method consistently outperforms the compared audio classification methods.
翻译:本文提出SemanticAC,一种语义辅助的音频分类框架,以更好地利用语义信息。与将类别标签视为离散向量的传统音频分类方法不同,我们采用语言模型从标签中提取丰富的语义信息,并优化音频信号与其标签之间的语义一致性。我们验证了仅利用标签的简单文本信息与先进预训练模型,即可提供更丰富的语义监督,从而提升性能。具体而言,我们设计了一个文本编码器来捕获标签文本扩展中的语义信息,随后通过音频编码器与相似度计算模块将音频信号映射至对应类别标签的语义空间,强制实现语义一致性。在ESC-50与US8K两个音频数据集上的大量实验表明,本方法在性能上始终优于对比的音频分类方法。