Retrieving adverbs that describe an action in a video poses a crucial step towards fine-grained video understanding. We propose a framework for video-to-adverb retrieval (and vice versa) that aligns video embeddings with their matching compositional adverb-action text embedding in a joint embedding space. The compositional adverb-action text embedding is learned using a residual gating mechanism, along with a novel training objective consisting of triplet losses and a regression target. Our method achieves state-of-the-art performance on five recent benchmarks for video-adverb retrieval. Furthermore, we introduce dataset splits to benchmark video-adverb retrieval for unseen adverb-action compositions on subsets of the MSR-VTT Adverbs and ActivityNet Adverbs datasets. Our proposed framework outperforms all prior works for the generalisation task of retrieving adverbs from videos for unseen adverb-action compositions. Code and dataset splits are available at https://hummelth.github.io/ReGaDa/.
翻译:检索描述视频中动作的副词是实现细粒度视频理解的关键步骤。我们提出了一种用于视频到副词检索(反之亦然)的框架,该框架在联合嵌入空间中将视频嵌入与其匹配的组合式副词-动作文本嵌入对齐。组合式副词-动作文本嵌入通过残差门控机制学习,并采用包含三元组损失和回归目标的新型训练目标。我们的方法在五个最新的视频-副词检索基准测试中达到了最先进性能。此外,我们引入了数据集划分,用于在MSR-VTT Adverbs和ActivityNet Adverbs数据集的子集上对未见过的副词-动作组合进行视频-副词检索基准测试。对于从未见过副词-动作组合的视频中检索副词的泛化任务,我们提出的框架优于所有先前工作。代码和数据集划分可在https://hummelth.github.io/ReGaDa/获取。