The rising popularity of multimodal large language models (MLLMs) has sparked a significant increase in research dedicated to evaluating these models. However, current evaluation studies predominantly concentrate on the ability of models to comprehend and reason within a unimodal (vision-only) context, overlooking critical performance evaluations in complex multimodal reasoning tasks that integrate both visual and text contexts. Furthermore, tasks that demand reasoning across multiple modalities pose greater challenges and require a deep understanding of multimodal contexts. In this paper, we introduce a comprehensive assessment framework named MM-InstructEval, which integrates a diverse array of metrics to provide an extensive evaluation of the performance of various models and instructions across a broad range of multimodal reasoning tasks with vision-text contexts. MM-InstructEval enhances the research on the performance of MLLMs in complex multimodal reasoning tasks, facilitating a more thorough and holistic zero-shot evaluation of MLLMs. We firstly utilize the "Best Performance" metric to determine the upper performance limit of each model across various datasets. The "Mean Relative Gain" metric provides an analysis of the overall performance across different models and instructions, while the "Stability" metric evaluates their sensitivity to variations. Historically, the research has focused on evaluating models independently or solely assessing instructions, overlooking the interplay between models and instructions. To address this gap, we introduce the "Adaptability" metric, designed to quantify the degree of adaptability between models and instructions. Evaluations are conducted on 31 models (23 MLLMs) across 16 multimodal datasets, covering 6 tasks, with 10 distinct instructions. The extensive analysis enables us to derive novel insights.
翻译:多模态大语言模型(MLLMs)的日益普及推动了大量专注于模型评估的研究。然而,当前的评估研究主要集中于模型在单模态(仅视觉)情境下理解和推理的能力,忽略了在融合视觉与文本背景的复杂多模态推理任务中的关键性能评估。此外,需要跨多个模态进行推理的任务构成了更大的挑战,要求对多模态上下文有深刻理解。本文提出一个名为MM-InstructEval的综合评估框架,该框架整合了多种指标,旨在广泛评估各类模型与指令在视觉-文本背景下的多模态推理任务中的表现。MM-InstructEval深化了对MLLMs在复杂多模态推理任务中性能的研究,推动了MLLMs更全面、整体的零样本评估。我们首先利用“最佳性能”指标来确定各模型在不同数据集上的性能上限。“平均相对增益”指标用于分析不同模型与指令的总体性能,而“稳定性”指标则评估其对变化的敏感度。以往研究侧重于独立评估模型或仅评估指令,忽视了模型与指令间的相互作用。为填补这一空白,我们引入“适应性”指标,旨在量化模型与指令之间的适配程度。我们对涵盖6类任务的16个多模态数据集上的31个模型(包括23个MLLMs)进行了评估,涉及10种不同的指令。广泛的分析使我们得以衍生出新颖的洞见。