Visual reasoning requires multimodal perception and commonsense cognition of the world. Recently, multiple vision-language models (VLMs) have been proposed with excellent commonsense reasoning ability in various domains. However, how to harness the collective power of these complementary VLMs is rarely explored. Existing methods like ensemble still struggle to aggregate these models with the desired higher-order communications. In this work, we propose Cola, a novel paradigm that coordinates multiple VLMs for visual reasoning. Our key insight is that a large language model (LLM) can efficiently coordinate multiple VLMs by facilitating natural language communication that leverages their distinct and complementary capabilities. Extensive experiments demonstrate that our instruction tuning variant, Cola-FT, achieves state-of-the-art performance on visual question answering (VQA), outside knowledge VQA, visual entailment, and visual spatial reasoning tasks. Moreover, we show that our in-context learning variant, Cola-Zero, exhibits competitive performance in zero and few-shot settings, without finetuning. Through systematic ablation studies and visualizations, we validate that a coordinator LLM indeed comprehends the instruction prompts as well as the separate functionalities of VLMs; it then coordinates them to enable impressive visual reasoning capabilities.
翻译:视觉推理需要多模态感知和对世界的常识认知。近年来,多种视觉语言模型(VLM)已被提出,在不同领域展现出卓越的常识推理能力。然而,如何整合这些互补的VLM的集体力量仍鲜有探索。现有的集成方法等仍然难以实现具有所需高阶通信的模型聚合。在本工作中,我们提出Cola,一种协调多个VLM进行视觉推理的新范式。我们的关键洞察是,大语言模型(LLM)可以通过促进利用它们独特且互补能力的自然语言通信,高效协调多个VLM。大量实验表明,我们的指令微调变体Cola-FT在视觉问答(VQA)、外部知识VQA、视觉蕴含和视觉空间推理任务上达到了最先进性能。此外,我们展示了上下文学习变体Cola-Zero在零样本和少样本设置下无需微调即展现出具有竞争力的性能。通过系统的消融研究和可视化,我们验证了协调者LLM确实能理解指令提示以及VLM的各自功能;随后它协调这些模型以实现令人印象深刻的视觉推理能力。