The surge of Multimodal Large Language Models (MLLMs), given their prominent emergent capabilities in instruction following and reasoning, has greatly advanced the field of visual reasoning. However, constrained by their non-lossless image tokenization, most MLLMs fall short of comprehensively capturing details of text and objects, especially in high-resolution images. To address this, we propose P2G, a novel framework for plug-and-play grounding of reasoning in MLLMs. Specifically, P2G exploits the tool-usage potential of MLLMs to employ expert agents to achieve on-the-fly grounding to critical visual and textual objects of image, thus achieving deliberate reasoning via multimodal prompting. We further create P2GB, a benchmark aimed at assessing MLLMs' ability to understand inter-object relationships and text in challenging high-resolution images. Comprehensive experiments on visual reasoning tasks demonstrate the superiority of P2G. Noteworthy, P2G achieved comparable performance with GPT-4V on P2GB, with a 7B backbone. Our work highlights the potential of plug-and-play grounding of reasoning and opens up a promising alternative beyond model scaling.
翻译:多模态大语言模型(MLLMs)因其在指令遵循与推理方面突出的涌现能力,极大地推动了视觉推理领域的发展。然而,受限于非无损图像分词化机制,大多数MLLMs难以全面捕捉文本和物体的细节,尤其是在高分辨率图像中。为解决这一问题,我们提出P2G——一种面向多模态大语言模型推理接地的即插即用框架。具体而言,P2G利用MLLMs的工具调用潜能,通过调用专家智能体实现对图像中关键视觉与文本实体的动态接地,并借助多模态提示实现深思熟虑的推理。我们进一步创建了P2GB基准测试,旨在评估MLLMs在挑战性高分辨率图像中理解物体间关系及文本的能力。在视觉推理任务上的全面实验证明了P2G的优越性。值得注意的是,基于7B骨干网络的P2G在P2GB上取得了与GPT-4V相当的性能。我们的工作凸显了即插即用推理接地的潜力,并为超越模型缩放提供了有前景的替代方案。