Visual and linguistic pre-training aims to learn vision and language representations together, which can be transferred to visual-linguistic downstream tasks. However, there exists semantic confusion between language and vision during the pre-training stage. Moreover, current pre-trained models tend to take lots of computation resources for fine-tuning when transferred to downstream tasks. In this work, we present a simple but effective approach for learning Contrastive and Adaptive representations of Vision and Language, namely CAVL. Specifically, we introduce a pair-wise contrastive loss to learn alignments between the whole sentence and each image in the same batch during the pre-training process. At the fine-tuning stage, we introduce two lightweight adaptation networks to reduce model parameters and increase training speed for saving computation resources. We evaluate our CAVL on six main downstream tasks, including Visual Question Answering (VQA), Visual Commonsense Reasoning (VCR), Natural Language for Visual Reasoning (NLVR), Region-to-Phrase Grounding (RPG), Text-to-Image Retrieval (TIR), and Zero-shot Text-to-Image Retrieval (ZS-TIR). Compared to baselines, we achieve superior performance and reduce the fine-tuning time by a large margin (in particular, 76.17%). Extensive experiments and ablation studies demonstrate the efficiency of contrastive pre-training and adaptive fine-tuning proposed in our CAVL.
翻译:视觉与语言预训练旨在共同学习视觉和语言表示,并将其迁移至视觉语言下游任务。然而,在预训练阶段,语言与视觉之间存在语义混淆。此外,当前预训练模型在迁移至下游任务时通常需要消耗大量计算资源进行微调。本文提出一种简单而有效的方法——CAVL,用于学习视觉与语言的对比式和自适应表示。具体而言,我们在预训练过程中引入成对对比损失,以学习同一批次内整个句子与每张图像之间的对齐关系。在微调阶段,我们引入两个轻量级自适应网络,以减少模型参数、提升训练速度,从而节省计算资源。我们在六个主要下游任务上评估CAVL,包括视觉问答(VQA)、视觉常识推理(VCR)、视觉推理的自然语言(NLVR)、区域到短语定位(RPG)、文本到图像检索(TIR)以及零样本文本到图像检索(ZS-TIR)。与基线相比,我们取得了更优性能,并大幅减少了微调时间(尤其降低了76.17%)。大量实验和消融研究表明,我们提出的CAVL中对比式预训练和自适应微调具有高效性。