In recent years, datasets of paired audio and captions have enabled remarkable success in automatically generating descriptions for audio clips, namely Automated Audio Captioning (AAC). However, it is labor-intensive and time-consuming to collect a sufficient number of paired audio and captions. Motivated by the recent advances in Contrastive Language-Audio Pretraining (CLAP), we propose a weakly-supervised approach to train an AAC model assuming only text data and a pre-trained CLAP model, alleviating the need for paired target data. Our approach leverages the similarity between audio and text embeddings in CLAP. During training, we learn to reconstruct the text from the CLAP text embedding, and during inference, we decode using the audio embeddings. To mitigate the modality gap between the audio and text embeddings we employ strategies to bridge the gap during training and inference stages. We evaluate our proposed method on Clotho and AudioCaps datasets demonstrating its ability to achieve a relative performance of up to ~$83\%$ compared to fully supervised approaches trained with paired target data.
翻译:近年来,配对音频与文字描述的数据集在自动生成音频片段描述(即自动音频描述,AAC)方面取得了显著成功。然而,收集足够数量的配对音频与文字描述数据既费时又费力。受对比语言-音频预训练(CLAP)最新进展的启发,我们提出了一种弱监督方法来训练AAC模型,该方法仅需文本数据和预训练的CLAP模型,从而减轻了对配对目标数据的依赖。我们的方法利用了音频和文本嵌入在CLAP中的相似性。在训练阶段,我们学习从CLAP文本嵌入重构文本;在推理阶段,则利用音频嵌入进行解码。为缓解音频与文本嵌入之间的模态差异,我们在训练和推理阶段采用了弥合这一差距的策略。我们在Clotho和AudioCaps数据集上评估了所提方法,结果表明其相对性能可达使用配对目标数据的全监督方法的约83%。