Text-video retrieval aims to find the most relevant cross-modal samples for a given query. Recent methods focus on modeling the whole spatial-temporal relations. However, since video clips contain more diverse content than captions, the model aligning these asymmetric video-text pairs has a high risk of retrieving many false positive results. In this paper, we propose Probabilistic Token Aggregation (\textit{ProTA}) to handle cross-modal interaction with content asymmetry. Specifically, we propose dual partial-related aggregation to disentangle and re-aggregate token representations in both low-dimension and high-dimension spaces. We propose token-based probabilistic alignment to generate token-level probabilistic representation and maintain the feature representation diversity. In addition, an adaptive contrastive loss is proposed to learn compact cross-modal distribution space. Based on extensive experiments, \textit{ProTA} achieves significant improvements on MSR-VTT (50.9%), LSMDC (25.8%), and DiDeMo (47.2%).
翻译:文本-视频检索旨在为给定查询找到最相关的跨模态样本。现有方法主要聚焦于建模全局时空关系。然而,由于视频片段包含比字幕更多样化的内容,对齐这些非对称视频-文本对的模型存在检索到大量假阳性结果的高风险。本文提出概率型词元聚合(ProTA)以处理存在内容不对称性的跨模态交互。具体而言,我们提出双部分相关聚合方法,在低维和高维空间中实现词元表征的解耦与重聚合。我们提出基于词元的概率对齐方法,生成词元级概率表征并保持特征表征多样性。此外,我们提出自适应对比损失以学习紧致的跨模态分布空间。基于大量实验,ProTA在MSR-VTT(50.9%)、LSMDC(25.8%)和DiDeMo(47.2%)数据集上均取得显著性能提升。