While recent progress in video-text retrieval has been advanced by the exploration of better representation learning, in this paper, we present a novel multi-space multi-grained supervised learning framework, SUMA, to learn an aligned representation space shared between the video and the text for video-text retrieval. The shared aligned space is initialized with a finite number of concept clusters, each of which refers to a number of basic concepts (words). With the text data at hand, we are able to update the shared aligned space in a supervised manner using the proposed similarity and alignment losses. Moreover, to enable multi-grained alignment, we incorporate frame representations for better modeling the video modality and calculating fine-grained and coarse-grained similarity. Benefiting from learned shared aligned space and multi-grained similarity, extensive experiments on several video-text retrieval benchmarks demonstrate the superiority of SUMA over existing methods.
翻译:尽管近期视频-文本检索领域的进展得益于对更优表示学习的探索,本文提出了一种新颖的多空间多粒度监督学习框架SUMA,用于学习视频与文本之间共享的对齐表示空间,以实现视频-文本检索。该共享对齐空间通过有限数量的概念簇进行初始化,每个概念簇对应一组基础概念(词语)。借助文本数据,我们能够利用所提出的相似度损失和对齐损失以监督方式更新共享对齐空间。此外,为实现多粒度对齐,我们引入帧表示以更好地建模视频模态,并计算细粒度与粗粒度相似度。得益于学习到的共享对齐空间与多粒度相似度,在多个视频-文本检索基准上的大量实验表明,SUMA方法显著优于现有方法。