Image-text contrastive models such as CLIP are useful for a variety of downstream applications including zero-shot classification, image-text retrieval and transfer learning. However, these contrastively trained vision-language models often fail on compositional visio-linguistic tasks such as Winoground with performance equivalent to random chance. In our paper, we address this issue and propose a sample-efficient light-weight method called SDS-CLIP to improve the compositional visio-linguistic reasoning capabilities of CLIP. The core idea of our method is to use differentiable image parameterizations to fine-tune CLIP with a distillation objective from large text-to-image generative models such as Stable-Diffusion which are relatively good at visio-linguistic reasoning tasks. On the challenging Winoground compositional reasoning benchmark, our method improves the absolute visio-linguistic performance of different CLIP models by up to 7%, while on the ARO dataset, our method improves the visio-linguistic performance by upto 3%. As a byproduct of inducing visio-linguistic reasoning into CLIP, we also find that the zero-shot performance improves marginally on a variety of downstream datasets. Our method reinforces that carefully designed distillation objectives from generative models can be leveraged to extend existing contrastive image-text models with improved visio-linguistic reasoning capabilities.
翻译:图像-文本对比模型(如CLIP)在多种下游应用中表现出色,包括零样本分类、图像-文本检索和迁移学习。然而,这些经过对比训练的视觉语言模型在组合性视觉语言任务(如Winoground)中常表现不佳,其性能与随机猜测相当。本文针对这一问题,提出一种轻量级、样本高效的方法SDS-CLIP,旨在提升CLIP的组合性视觉语言推理能力。该方法的核心思想是利用可微图像参数化,通过从大型文本到图像生成模型(如Stable-Diffusion,这类模型在视觉语言推理任务中表现相对较好)中提取蒸馏目标,对CLIP进行微调。在具有挑战性的Winoground组合推理基准测试中,我们的方法将不同CLIP模型的绝对视觉语言性能提升了高达7%;在ARO数据集上,视觉语言性能提升了高达3%。作为在CLIP中引入视觉语言推理能力的附带结果,我们还发现其在多种下游数据集上的零样本性能略有提升。我们的方法表明,精心设计的生成模型蒸馏目标可有效扩展现有对比性图像-文本模型,增强其视觉语言推理能力。