Understanding verbs is crucial to modelling how people and objects interact with each other and the environment through space and time. Recently, state-of-the-art video-language models based on CLIP have been shown to have limited verb understanding and to rely extensively on nouns, restricting their performance in real-world video applications that require action and temporal understanding. In this work, we improve verb understanding for CLIP-based video-language models by proposing a new Verb-Focused Contrastive (VFC) framework. This consists of two main components: (1) leveraging pretrained large language models (LLMs) to create hard negatives for cross-modal contrastive learning, together with a calibration strategy to balance the occurrence of concepts in positive and negative pairs; and (2) enforcing a fine-grained, verb phrase alignment loss. Our method achieves state-of-the-art results for zero-shot performance on three downstream tasks that focus on verb understanding: video-text matching, video question-answering and video classification. To the best of our knowledge, this is the first work which proposes a method to alleviate the verb understanding problem, and does not simply highlight it.
翻译:理解动词对于建模人与物以及环境在时空中的交互至关重要。近期,基于CLIP的最先进的视频-语言模型被证明在理解动词方面存在局限,且严重依赖名词,这限制了它们在需要动作和时序理解的真实视频应用中的表现。在本研究中,我们通过提出一种新的动词聚焦对比框架(Verb-Focused Contrastive, VFC)来改进基于CLIP的视频-语言模型中的动词理解。该框架包含两个主要组成部分:(1)利用预训练的大语言模型(LLMs)生成跨模态对比学习中的困难负样本,并采用校准策略来平衡正负样本对中概念的分布;(2)强制执行细粒度的动词短语对齐损失。我们的方法在三个聚焦动词理解的下游任务(视频-文本匹配、视频问答和视频分类)的零样本性能上达到了最先进水平。据我们所知,这是首个提出方法以缓解动词理解问题的工作,而并非仅指出该问题。