Video-language embeddings are a promising avenue for injecting semantics into visual representations, but existing methods capture only short-term associations between seconds-long video clips and their accompanying text. We propose HierVL, a novel hierarchical video-language embedding that simultaneously accounts for both long-term and short-term associations. As training data, we take videos accompanied by timestamped text descriptions of human actions, together with a high-level text summary of the activity throughout the long video (as are available in Ego4D). We introduce a hierarchical contrastive training objective that encourages text-visual alignment at both the clip level and video level. While the clip-level constraints use the step-by-step descriptions to capture what is happening in that instant, the video-level constraints use the summary text to capture why it is happening, i.e., the broader context for the activity and the intent of the actor. Our hierarchical scheme yields a clip representation that outperforms its single-level counterpart as well as a long-term video representation that achieves SotA results on tasks requiring long-term video modeling. HierVL successfully transfers to multiple challenging downstream tasks (in EPIC-KITCHENS-100, Charades-Ego, HowTo100M) in both zero-shot and fine-tuned settings.
翻译:视频-语言嵌入为视觉表示注入语义提供了有前景的途径,但现有方法仅能捕捉秒级视频片段与其对应文本之间的短期关联。我们提出HierVL——一种新颖的分层视频-语言嵌入方法,可同时建模长期与短期关联。训练数据采用带时间戳的人类动作文本描述的长视频,以及覆盖整个视频活动的高层文本摘要(如Ego4D数据集中所提供的)。我们引入层级对比训练目标,在片段级和视频级同时促进文本-视觉对齐:片段级约束利用逐步描述捕捉瞬时事件,视频级约束则通过摘要文本揭示事件发生的原因(即活动的更广泛背景与行为者意图)。这种分层方案生成的片段表示优于单层对应方法,同时其长期视频表示在需要长期视频建模的任务中达到SotA性能。HierVL在零样本和微调设置下均能成功迁移至多个具有挑战性的下游任务(EPIC-KITCHENS-100、Charades-Ego、HowTo100M)。