Finding the right sound effects (SFX) to match moments in a video is a difficult and time-consuming task, and relies heavily on the quality and completeness of text metadata. Retrieving high-quality (HQ) SFX using a video frame directly as the query is an attractive alternative, removing the reliance on text metadata and providing a low barrier to entry for non-experts. Due to the lack of HQ audio-visual training data, previous work on audio-visual retrieval relies on YouTube (in-the-wild) videos of varied quality for training, where the audio is often noisy and the video of amateur quality. As such it is unclear whether these systems would generalize to the task of matching HQ audio to production-quality video. To address this, we propose a multimodal framework for recommending HQ SFX given a video frame by (1) leveraging large language models and foundational vision-language models to bridge HQ audio and video to create audio-visual pairs, resulting in a highly scalable automatic audio-visual data curation pipeline; and (2) using pre-trained audio and visual encoders to train a contrastive learning-based retrieval system. We show that our system, trained using our automatic data curation pipeline, significantly outperforms baselines trained on in-the-wild data on the task of HQ SFX retrieval for video. Furthermore, while the baselines fail to generalize to this task, our system generalizes well from clean to in-the-wild data, outperforming the baselines on a dataset of YouTube videos despite only being trained on the HQ audio-visual pairs. A user study confirms that people prefer SFX retrieved by our system over the baseline 67% of the time both for HQ and in-the-wild data. Finally, we present ablations to determine the impact of model and data pipeline design choices on downstream retrieval performance. Please visit our project website to listen to and view our SFX retrieval results.
翻译:为视频中的时刻匹配恰当的音效(SFX)是一项困难且耗时的任务,且高度依赖文本元数据的质量和完整性。直接使用视频帧作为查询来检索高质量(HQ)音效是一种极具吸引力的替代方案,它消除了对文本元数据的依赖,并为非专业人士提供了低门槛的接入方式。由于缺乏高质量音视频训练数据,以往的音视频检索研究依赖于质量参差不齐的YouTube(野外)视频进行训练,其音频常有噪声且视频为业余质量。因此,尚不清楚此类系统是否能泛化到将高质量音频与专业质量视频匹配的任务。为解决这一问题,我们提出一种多模态框架,通过以下方式为给定视频帧推荐高质量音效:(1)利用大语言模型和基础视觉-语言模型桥接高质量音频与视频,创建音视频对,形成高度可扩展的自动音视频数据整理管道;(2)使用预训练的音频和视觉编码器训练基于对比学习的检索系统。实验表明,采用自动数据整理管道训练的系统在视频高质量音效检索任务上显著优于基于野外数据训练的基线方法。此外,尽管基线方法无法泛化至该任务,我们的系统却能很好地从干净数据泛化到野外数据,在YouTube视频数据集上的表现优于基线方法,即便其仅使用高质量音视频对进行训练。用户研究证实,无论是高质量数据还是野外数据,用户有67%的概率更偏好我们系统检索的音效而非基线方法。最终,我们通过消融实验确定了模型与数据管道设计选择对下游检索性能的影响。欢迎访问项目网站聆听和观看我们的音效检索结果。