The advancement of audio-language (AL) multimodal learning tasks has been significant in recent years. However, researchers face challenges due to the costly and time-consuming collection process of existing audio-language datasets, which are limited in size. To address this data scarcity issue, we introduce WavCaps, the first large-scale weakly-labelled audio captioning dataset, comprising approximately 400k audio clips with paired captions. We sourced audio clips and their raw descriptions from web sources and a sound event detection dataset. However, the online-harvested raw descriptions are highly noisy and unsuitable for direct use in tasks such as automated audio captioning. To overcome this issue, we propose a three-stage processing pipeline for filtering noisy data and generating high-quality captions, where ChatGPT, a large language model, is leveraged to filter and transform raw descriptions automatically. We conduct a comprehensive analysis of the characteristics of WavCaps dataset and evaluate it on multiple downstream audio-language multimodal learning tasks. The systems trained on WavCaps outperform previous state-of-the-art (SOTA) models by a significant margin. Our aspiration is for the WavCaps dataset we have proposed to facilitate research in audio-language multimodal learning and demonstrate the potential of utilizing ChatGPT to enhance academic research. Our dataset and codes are available at https://github.com/XinhaoMei/WavCaps.
翻译:近年来,音频-语言多模态学习任务取得了显著进展。然而,由于现有音频-语言数据集的收集过程成本高昂且耗时,且规模有限,研究者面临诸多挑战。为应对数据稀缺问题,我们提出了WavCaps——首个大规模弱标签音频字幕数据集,包含约40万个带配对字幕的音频片段。我们从网络资源和一个声音事件检测数据集中获取了音频片段及其原始描述。然而,这些在线收集的原始描述含有大量噪声,无法直接用于自动音频字幕生成等任务。为解决这一问题,我们提出了一个三阶段处理流程,用于过滤噪声数据并生成高质量字幕,其中利用了大型语言模型ChatGPT自动过滤和转换原始描述。我们对WavCaps数据集的特征进行了全面分析,并在多个下游音频-语言多模态学习任务上对其进行了评估。基于WavCaps训练的系统显著超越了以往的最先进模型。我们期望所提出的WavCaps数据集能够促进音频-语言多模态学习的研究,并展示利用ChatGPT增强学术研究的潜力。我们的数据集和代码公开于https://github.com/XinhaoMei/WavCaps。