This paper considers contrastive training for cross-modal 0-shot transfer wherein a pre-trained model in one modality is used for representation learning in another domain using pairwise data. The learnt models in the latter domain can then be used for a diverse set of tasks in a zero-shot way, similar to ``Contrastive Language-Image Pre-training (CLIP)'' and ``Locked-image Tuning (LiT)'' that have recently gained considerable attention. Most existing works for cross-modal representation alignment (including CLIP and LiT) use the standard contrastive training objective, which employs sets of positive and negative examples to align similar and repel dissimilar training data samples. However, similarity amongst training examples has a more continuous nature, thus calling for a more `non-binary' treatment. To address this, we propose a novel loss function called Continuously Weighted Contrastive Loss (CWCL) that employs a continuous measure of similarity. With CWCL, we seek to align the embedding space of one modality with another. Owing to the continuous nature of similarity in the proposed loss function, these models outperform existing methods for 0-shot transfer across multiple models, datasets and modalities. Particularly, we consider the modality pairs of image-text and speech-text and our models achieve 5-8% (absolute) improvement over previous state-of-the-art methods in 0-shot image classification and 20-30% (absolute) improvement in 0-shot speech-to-intent classification and keyword classification.
翻译:本文研究跨模态零样本迁移中的对比训练方法,即利用成对数据将一个模态的预训练模型迁移至另一领域进行表征学习。此类习得的模型可像近期备受关注的"对比语言-图像预训练(CLIP)"和"锁定图像微调(LiT)"方法一样,以零样本方式完成多样化的下游任务。现有跨模态表征对齐工作(包括CLIP和LiT)大多采用标准对比训练目标函数,通过正负样本集对齐相似样本并分离非相似训练样本。然而,训练样本间的相似性具有连续分布特性,需要更"非二元"的处理方式。为此,我们提出名为连续加权对比损失(CWCL)的新型损失函数,该函数采用连续相似性度量。通过CWCL,我们旨在将一个模态的嵌入空间与另一模态对齐。得益于所提损失函数中相似性的连续特性,这些模型在多种模型、数据集和模态的零样本迁移任务中均优于现有方法。特别地,我们针对图像-文本和语音-文本模态对进行实验,所提模型在零样本图像分类任务中相比先前最优方法提升5-8%(绝对值),在零样本语音意图分类和关键词分类任务中提升20-30%(绝对值)。