Video databases from the internet are a valuable source of text-audio retrieval datasets. However, given that sound and vision streams represent different "views" of the data, treating visual descriptions as audio descriptions is far from optimal. Even if audio class labels are present, they commonly are not very detailed, making them unsuited for text-audio retrieval. To exploit relevant audio information from video-text datasets, we introduce a methodology for generating audio-centric descriptions using Large Language Models (LLMs). In this work, we consider the egocentric video setting and propose three new text-audio retrieval benchmarks based on the EpicMIR and EgoMCQ tasks, and on the EpicSounds dataset. Our approach for obtaining audio-centric descriptions gives significantly higher zero-shot performance than using the original visual-centric descriptions. Furthermore, we show that using the same prompts, we can successfully employ LLMs to improve the retrieval on EpicSounds, compared to using the original audio class labels of the dataset. Finally, we confirm that LLMs can be used to determine the difficulty of identifying the action associated with a sound.
翻译:来自互联网的视频数据库是文本-音频检索数据集的重要来源。然而,鉴于声音和视觉流代表数据的不同“视角”,将视觉描述视为音频描述远非最优。即使存在音频类别标签,它们通常也不够详细,因此不适合文本-音频检索。为了利用视频文本数据集中的相关音频信息,我们引入了一种使用大型语言模型(LLMs)生成以音频为中心的描述的方法。在本工作中,我们考虑自我中心视频设置,并基于EpicMIR和EgoMCQ任务以及EpicSounds数据集提出三个新的文本-音频检索基准。与使用原始视觉描述相比,我们获取以音频为中心描述的方法在零样本性能上显著提升。此外,我们表明,使用相同的提示,与使用数据集的原始音频类别标签相比,我们可以成功利用LLMs改进EpicSounds上的检索性能。最后,我们确认LLMs可用于确定与声音相关的动作识别的难度。