Text-to-video retrieval systems have recently made significant progress by utilizing pre-trained models trained on large-scale image-text pairs. However, most of the latest methods primarily focus on the video modality while disregarding the audio signal for this task. Nevertheless, a recent advancement by ECLIPSE has improved long-range text-to-video retrieval by developing an audiovisual video representation. Nonetheless, the objective of the text-to-video retrieval task is to capture the complementary audio and video information that is pertinent to the text query rather than simply achieving better audio and video alignment. To address this issue, we introduce TEFAL, a TExt-conditioned Feature ALignment method that produces both audio and video representations conditioned on the text query. Instead of using only an audiovisual attention block, which could suppress the audio information relevant to the text query, our approach employs two independent cross-modal attention blocks that enable the text to attend to the audio and video representations separately. Our proposed method's efficacy is demonstrated on four benchmark datasets that include audio: MSR-VTT, LSMDC, VATEX, and Charades, and achieves better than state-of-the-art performance consistently across the four datasets. This is attributed to the additional text-query-conditioned audio representation and the complementary information it adds to the text-query-conditioned video representation.
翻译:文本-视频检索系统近年来通过利用在大规模图像-文本对上预训练的模型取得了显著进展。然而,大多数最新方法主要聚焦于视频模态,而忽略了该任务中的音频信号。尽管如此,ECLIPSE的最新进展通过开发音视频联合表示改进了长程文本-视频检索。但文本-视频检索任务的目标是捕获与文本查询相关的互补性音频和视频信息,而不仅仅是实现更好的音频与视频对齐。针对这一问题,我们提出了TEFAL——一种基于文本条件特征对齐方法,该方法生成与文本查询条件相关的音频和视频表示。我们的方法不采用可能抑制文本查询相关音频信息的单模块音视频注意力机制,而是使用两个独立的跨模态注意力块,使文本能够分别关注音频和视频表示。我们在包含音频的四个基准数据集(MSR-VTT、LSMDC、VATEX和Charades)上验证了该方法的效果,并在所有四个数据集上持续取得优于现有最优方法的性能。这一成果归因于额外的文本查询条件音频表示及其为文本查询条件视频表示补充的互补信息。