With the explosion of multimedia content in recent years, Video Corpus Moment Retrieval (VCMR), which aims to detect a video moment that matches a given natural language query from multiple videos, has become a critical problem. However, existing VCMR studies have a significant limitation since they have regarded all videos not paired with a specific query as negative, neglecting the possibility of including false negatives when constructing the negative video set. In this paper, we propose an MVMR (Massive Videos Moment Retrieval) task that aims to localize video frames within a massive video set, mitigating the possibility of falsely distinguishing positive and negative videos. For this task, we suggest an automatic dataset construction framework by employing textual and visual semantic matching evaluation methods on the existing video moment search datasets and introduce three MVMR datasets. To solve MVMR task, we further propose a strong method, CroCs, which employs cross-directional contrastive learning that selectively identifies the reliable and informative negatives, enhancing the robustness of a model on MVMR task. Experimental results on the introduced datasets reveal that existing video moment search models are easily distracted by negative video frames, whereas our model shows significant performance.
翻译:近年来,随着多媒体内容的爆炸式增长,视频语料库时刻检索(VCMR)成为一个关键问题,其旨在从多个视频中检测与给定自然语言查询匹配的视频片段。然而,现有VCMR研究存在一个重大局限,即将所有未与特定查询配对的视频视为负样本,忽略了在构建负视频集时可能包含假负例的情况。本文提出了一项MVMR(海量视频时刻检索)任务,旨在在海量视频集中定位视频帧,从而减少错误区分正负视频的可能性。针对该任务,我们提出了一种自动数据集构建框架,通过利用现有视频时刻搜索数据集上的文本与视觉语义匹配评估方法,并引入了三个MVMR数据集。为解决MVMR任务,我们进一步提出了一种强效方法CroCs,该方法采用跨方向对比学习,有选择地识别可靠且信息丰富的负样本,从而增强模型在MVMR任务上的鲁棒性。在引入的数据集上的实验结果表明,现有视频时刻搜索模型易受负视频帧干扰,而我们的模型表现出显著性能提升。