Contrastive language-image pre-training (CLIP) models have demonstrated considerable success across various vision-language tasks, such as text-to-image retrieval, where the model is required to effectively process natural language input to produce an accurate visual output. However, current models still face limitations in dealing with linguistic variations in input queries, such as paraphrases, making it challenging to handle a broad range of user queries in real-world applications. In this study, we introduce a straightforward fine-tuning approach to enhance the representations of CLIP models for paraphrases. Our approach involves a two-step paraphrase generation process, where we automatically create two categories of paraphrases from web-scale image captions by leveraging large language models. Subsequently, we fine-tune the CLIP text encoder using these generated paraphrases while freezing the image encoder. Our resulting model, which we call ParaCLIP, exhibits significant improvements over baseline CLIP models across various tasks, including paraphrased retrieval (with rank similarity scores improved by up to 2.0% and 5.6%), Visual Genome Relation and Attribution, as well as seven semantic textual similarity tasks.
翻译:对比语言-图像预训练(CLIP)模型在各种视觉-语言任务中取得了显著成功,例如文本到图像检索,该任务要求模型有效处理自然语言输入以生成准确的视觉输出。然而,当前模型在处理输入查询的语言变化(如释义)方面仍存在局限性,这使得在现实应用中应对广泛的用户查询面临挑战。本研究提出了一种简单的微调方法,以增强CLIP模型对释义的表示能力。我们的方法涉及两步释义生成过程,通过利用大型语言模型从网络规模的图像标题中自动创建两类释义。随后,我们冻结图像编码器,使用生成的释义微调CLIP文本编码器。我们得到的模型称为ParaCLIP,在多种任务上相较于基线CLIP模型展现出显著改进,包括释义检索(排序相似度得分分别提升2.0%和5.6%)、Visual Genome关系与属性任务,以及七项语义文本相似性任务。