Multimodal Large Language Model (MLLM) relies on the powerful LLM to perform multimodal tasks, showing amazing emergent abilities in recent studies, such as writing poems based on an image. However, it is difficult for these case studies to fully reflect the performance of MLLM, lacking a comprehensive evaluation. In this paper, we fill in this blank, presenting the first comprehensive MLLM Evaluation benchmark MME. It measures both perception and cognition abilities on a total of 14 subtasks. In order to avoid data leakage that may arise from direct use of public datasets for evaluation, the annotations of instruction-answer pairs are all manually designed. The concise instruction design allows us to fairly compare MLLMs, instead of struggling in prompt engineering. Besides, with such an instruction, we can also easily carry out quantitative statistics. A total of 30 advanced MLLMs are comprehensively evaluated on our MME, which not only suggests that existing MLLMs still have a large room for improvement, but also reveals the potential directions for the subsequent model optimization.
翻译:多模态大型语言模型(MLLM)依赖强大的LLM执行多模态任务,近年来展现出惊人的涌现能力,例如根据图像创作诗歌。然而,这些案例研究难以充分反映MLLM的性能,缺乏全面评估。本文填补了这一空白,提出了首个综合性MLLM评估基准MME。该基准在总计14项子任务上同时测量感知与认知能力。为避免直接使用公开数据集进行评估可能导致的数据泄露,所有指令-答案对的标注均为人工设计。简洁的指令设计使我们能够公平比较不同MLLM,无需纠结于提示工程。此外,借助此类指令,我们还能便捷地进行量化统计。我们在MME上全面评估了30个先进MLLM,结果不仅表明现有MLLM仍有巨大改进空间,还揭示了后续模型优化的潜在方向。