Large-scale pre-trained Vision & Language (VL) models have shown remarkable performance in many applications, enabling replacing a fixed set of supported classes with zero-shot open vocabulary reasoning over (almost arbitrary) natural language prompts. However, recent works have uncovered a fundamental weakness of these models. For example, their difficulty to understand Visual Language Concepts (VLC) that go 'beyond nouns' such as the meaning of non-object words (e.g., attributes, actions, relations, states, etc.), or difficulty in performing compositional reasoning such as understanding the significance of the order of the words in a sentence. In this work, we investigate to which extent purely synthetic data could be leveraged to teach these models to overcome such shortcomings without compromising their zero-shot capabilities. We contribute Synthetic Visual Concepts (SyViC) - a million-scale synthetic dataset and data generation codebase allowing to generate additional suitable data to improve VLC understanding and compositional reasoning of VL models. Additionally, we propose a general VL finetuning strategy for effectively leveraging SyViC towards achieving these improvements. Our extensive experiments and ablations on VL-Checklist, Winoground, and ARO benchmarks demonstrate that it is possible to adapt strong pre-trained VL models with synthetic data significantly enhancing their VLC understanding (e.g. by 9.9% on ARO and 4.3% on VL-Checklist) with under 1% drop in their zero-shot accuracy.
翻译:大规模预训练的视觉与语言(VL)模型在众多应用中展现出卓越性能,能够将固定支持类别集替换为基于(几乎任意)自然语言提示的零样本开放词汇推理。然而,近期研究揭示了这些模型的根本弱点。例如,它们难以理解“超越名词”的视觉语言概念(VLC),如非物体词汇(属性、动作、关系、状态等)的含义,或在组合推理中遇到困难,例如无法理解句子中词语顺序的重要性。本研究探讨了纯合成数据在多大程度上可用于教导这些模型克服此类不足,同时不损害其零样本能力。我们贡献了合成视觉概念(SyViC)——一个百万规模的合成数据集及数据生成代码库,可生成额外合适数据以提升VL模型的VLC理解和组合推理能力。此外,我们提出了一种通用的VL微调策略,以有效利用SyViC实现这些改进。在VL-Checklist、Winoground和ARO基准上的广泛实验与消融研究表明,使用合成数据适配强预训练VL模型是可行的,能显著增强其VLC理解(例如,在ARO上提升9.9%,在VL-Checklist上提升4.3%),同时零样本准确率下降不足1%。