Large Multimodal Models (LMMs) such as GPT-4V and LLaVA have shown remarkable capabilities in visual reasoning with common image styles. However, their robustness against diverse style shifts, crucial for practical applications, remains largely unexplored. In this paper, we propose a new benchmark, BenchLMM, to assess the robustness of LMMs against three different styles: artistic image style, imaging sensor style, and application style, where each style has five sub-styles. Utilizing BenchLMM, we comprehensively evaluate state-of-the-art LMMs and reveal: 1) LMMs generally suffer performance degradation when working with other styles; 2) An LMM performs better than another model in common style does not guarantee its superior performance in other styles; 3) LMMs' reasoning capability can be enhanced by prompting LMMs to predict the style first, based on which we propose a versatile and training-free method for improving LMMs; 4) An intelligent LMM is expected to interpret the causes of its errors when facing stylistic variations. We hope that our benchmark and analysis can shed new light on developing more intelligent and versatile LMMs.
翻译:大型多模态模型(如GPT-4V和LLaVA)在处理常见图像风格的视觉推理任务中展现出卓越能力。然而,这些模型对不同风格变换的鲁棒性——这一对实际应用至关重要的特性——仍鲜有研究。本文提出新基准BenchLMM,用于评估LMMs在三种不同风格下的鲁棒性:艺术图像风格、成像传感器风格和应用风格,每种风格包含五个子类。借助BenchLMM,我们全面评估了当前最先进的LMMs,发现:1)LMMs处理非通用风格时普遍出现性能下降;2)在通用风格中表现更优的LMM,并不保证在其他风格中具有更优性能;3)通过引导LMMs先预测风格可增强其推理能力,并据此提出一种无需训练的通用改进方法;4)智能LMM应能解释面对风格变化时产生错误的原因。我们期望该基准与分析能为开发更智能、更通用的LMMs提供新思路。