Video grounding aims to locate a moment of interest matching the given query sentence from an untrimmed video. Previous works ignore the \emph{sparsity dilemma} in video annotations, which fails to provide the context information between potential events and query sentences in the dataset. In this paper, we contend that exploiting easily available captions which describe general actions \ie, prompt captions (PC) defined in our paper, will significantly boost the performance. To this end, we propose a Prompt Caption Network (PCNet) for video grounding. Specifically, we first introduce dense video captioning to generate dense captions and then obtain prompt captions by Non-Prompt Caption Suppression (NPCS). To capture the potential information in prompt captions, we propose Caption Guided Attention (CGA) project the semantic relations between prompt captions and query sentences into temporal space and fuse them into visual representations. Considering the gap between prompt captions and ground truth, we propose Asymmetric Cross-modal Contrastive Learning (ACCL) for constructing more negative pairs to maximize cross-modal mutual information. Without bells and whistles, extensive experiments on three public datasets (\ie, ActivityNet Captions, TACoS and ActivityNet-CG) demonstrate that our method significantly outperforms state-of-the-art methods.
翻译:视频定位旨在从非裁剪视频中定位与给定查询语句匹配的感兴趣时间段。以往的工作忽略了视频标注中的稀疏性困境,未能提供数据集中潜在事件与查询语句之间的上下文信息。在本文中,我们认为利用易于获取的描述一般动作的字幕(即本文定义的提示字幕)将显著提升性能。为此,我们提出了一种用于视频定位的提示字幕网络(PCNet)。具体而言,我们首先引入密集视频字幕生成密集字幕,然后通过非提示字幕抑制(NPCS)获得提示字幕。为了捕获提示字幕中的潜在信息,我们提出字幕引导注意力(CGA)将提示字幕与查询语句之间的语义关系映射到时域空间,并将其融合到视觉表示中。考虑到提示字幕与真实标注之间的差距,我们提出了非对称跨模态对比学习(ACCL)来构建更多负样本对,以最大化跨模态互信息。无需过多修饰,在三个公开数据集(即ActivityNet Captions、TACoS和ActivityNet-CG)上的广泛实验表明,我们的方法显著优于现有最先进方法。