Despite the impressive performance achieved by pre-trained language-and-vision models in downstream tasks, it remains an open question whether this reflects a proper understanding of image-text interaction. In this work, we explore to what extent they handle basic linguistic constructions -- active-passive voice, coordination, and relative clauses -- that even preschool children can typically master. We present BLA, a novel, automatically constructed benchmark to evaluate multimodal models on these Basic Language Abilities. We show that different types of Transformer-based systems, such as CLIP, ViLBERT, and BLIP2, generally struggle with BLA in a zero-shot setting, in line with previous findings. Our experiments, in particular, show that most of the tested models only marginally benefit when fine-tuned or prompted with construction-specific samples. Yet, the generative BLIP2 shows promising trends, especially in an in-context learning setting. This opens the door to using BLA not only as an evaluation benchmark but also to improve models' basic language abilities.
翻译:尽管预训练的语言-视觉模型在下游任务中取得了令人瞩目的性能,但这些表现是否真正反映了对图像-文本交互的恰当理解仍是一个开放问题。本研究探索了这些模型在多大程度上能处理儿童通常掌握的三种基础语言结构——主动-被动语态、并列结构和关系从句。我们提出了BLA,一个新型自动构建的基准测试,用于评估多模态模型在这些基础语言能力上的表现。研究表明,不同类型的基于Transformer的系统(如CLIP、ViLBERT和BLIP2)在零样本设置下普遍难以应对BLA任务,这与先前发现一致。我们的实验特别指出,当通过微调或提示学习使用特定结构样本时,大多数被测试模型仅获得有限改进。然而,生成式模型BLIP2展现出有前景的发展趋势,尤其在上下文学习场景中。这为将BLA不仅用作评估基准,还用于提升模型基础语言能力打开了新途径。