The introduction of audio latent diffusion models possessing the ability to generate realistic sound clips on demand from a text description has the potential to revolutionize how we work with audio. In this work, we make an initial attempt at understanding the inner workings of audio latent diffusion models by investigating how their audio outputs compare with the training data, similar to how a doctor auscultates a patient by listening to the sounds of their organs. Using text-to-audio latent diffusion models trained on the AudioCaps dataset, we systematically analyze memorization behavior as a function of training set size. We also evaluate different retrieval metrics for evidence of training data memorization, finding the similarity between mel spectrograms to be more robust in detecting matches than learned embedding vectors. In the process of analyzing memorization in audio latent diffusion models, we also discover a large amount of duplicated audio clips within the AudioCaps database.
翻译:具备从文本描述按需生成逼真声音片段的音频潜在扩散模型的引入,有可能彻底改变我们处理音频的方式。本研究通过探究音频输出与训练数据的比较方式,类似于医生通过听诊器官声音来检查患者,首次尝试理解音频潜在扩散模型的内部工作原理。利用在AudioCaps数据集上训练的文本到音频潜在扩散模型,我们系统分析了记忆行为随训练集规模的变化。我们还评估了不同检索指标以寻找训练数据记忆的证据,发现梅尔频谱图之间的相似性在检测匹配方面比学习嵌入向量更鲁棒。在分析音频潜在扩散模型记忆行为的过程中,我们还发现AudioCaps数据库中存在大量重复的音频片段。