Joint visual and language modeling on large-scale datasets has recently shown good progress in multi-modal tasks when compared to single modal learning. However, robustness of these approaches against real-world perturbations has not been studied. In this work, we perform the first extensive robustness study of video-language models against various real-world perturbations. We focus on text-to-video retrieval and propose two large-scale benchmark datasets, MSRVTT-P and YouCook2-P, which utilize 90 different visual and 35 different text perturbations. The study reveals some interesting initial findings from the studied models: 1) models are generally more susceptible when only video is perturbed as opposed to when only text is perturbed, 2) models that are pre-trained are more robust than those trained from scratch, 3) models attend more to scene and objects rather than motion and action. We hope this study will serve as a benchmark and guide future research in robust video-language learning. The benchmark introduced in this study along with the code and datasets is available at https://bit.ly/3CNOly4.
翻译:针对大规模数据集的联合视觉与语言建模近期在多模态任务中相较于单模态学习取得了显著进展,但这些方法在真实世界扰动下的鲁棒性尚未得到充分研究。本文首次系统研究了视频语言模型面对多种真实扰动的鲁棒性表现。我们聚焦于文本到视频检索任务,并构建了两个大规模基准数据集MSRVTT-P和YouCook2-P,分别包含90种视觉扰动和35种文本扰动。研究揭示了以下重要发现:1) 相较于仅对文本施加扰动,模型在仅对视频施加扰动时普遍更不稳定;2) 经过预训练的模型比从零训练的模型具有更强的鲁棒性;3) 模型更关注场景与物体特征,而非运动与动作特征。我们期望本研究能作为基准测试,为鲁棒视频语言学习的未来研究提供指导。本研究所提出的基准、代码及数据集已开源至https://bit.ly/3CNOly4。