Recent years have seen progress beyond domain-specific sound separation for speech or music towards universal sound separation for arbitrary sounds. Prior work on universal sound separation has investigated separating a target sound out of an audio mixture given a text query. Such text-queried sound separation systems provide a natural and scalable interface for specifying arbitrary target sounds. However, supervised text-queried sound separation systems require costly labeled audio-text pairs for training. Moreover, the audio provided in existing datasets is often recorded in a controlled environment, causing a considerable generalization gap to noisy audio in the wild. In this work, we aim to approach text-queried universal sound separation by using only unlabeled data. We propose to leverage the visual modality as a bridge to learn the desired audio-textual correspondence. The proposed CLIPSep model first encodes the input query into a query vector using the contrastive language-image pretraining (CLIP) model, and the query vector is then used to condition an audio separation model to separate out the target sound. While the model is trained on image-audio pairs extracted from unlabeled videos, at test time we can instead query the model with text inputs in a zero-shot setting, thanks to the joint language-image embedding learned by the CLIP model. Further, videos in the wild often contain off-screen sounds and background noise that may hinder the model from learning the desired audio-textual correspondence. To address this problem, we further propose an approach called noise invariant training for training a query-based sound separation model on noisy data. Experimental results show that the proposed models successfully learn text-queried universal sound separation using only noisy unlabeled videos, even achieving competitive performance against a supervised model in some settings.
翻译:近年来,从针对语音或音乐的特定领域声音分离,发展到针对任意声音的通用声音分离。先前关于通用声音分离的工作探索了根据文本查询从音频混合中分离出目标声音。此类文本查询声音分离系统为指定任意目标声音提供了自然且可扩展的接口。然而,有监督的文本查询声音分离系统需要昂贵的标注音频-文本对进行训练。此外,现有数据集提供的音频通常在受控环境中录制,导致其与野外含噪音频之间存在显著的泛化差距。本研究旨在仅使用无标签数据实现文本查询通用声音分离。我们提出利用视觉模态作为桥梁来学习所需的音频-文本对应关系。所提出的CLIPSep模型首先使用对比语言-图像预训练(CLIP)模型将输入查询编码为查询向量,随后该查询向量用于条件化音频分离模型以分离出目标声音。尽管模型基于从无标签视频中提取的图像-音频对进行训练,但在测试阶段,得益于CLIP模型学习到的联合语言-图像嵌入,我们可以通过零样本方式以文本输入查询模型。此外,野外视频常包含画外音和背景噪声,这可能阻碍模型学习所需的音频-文本对应关系。为解决此问题,我们进一步提出一种称为噪声不变训练的方法,用于在含噪数据上训练基于查询的声音分离模型。实验结果表明,所提出的模型仅使用含噪无标签视频即可成功学习文本查询通用声音分离,甚至在部分设置下达到了与有监督模型相当的性能。