The rapid evolution of Multi-modality Large Language Models (MLLMs) has catalyzed a shift in computer vision from specialized models to general-purpose foundation models. Nevertheless, there is still an inadequacy in assessing the abilities of MLLMs on low-level visual perception and understanding. To address this gap, we present Q-Bench, a holistic benchmark crafted to systematically evaluate potential abilities of MLLMs on three realms: low-level visual perception, low-level visual description, and overall visual quality assessment. a) To evaluate the low-level perception ability, we construct the LLVisionQA dataset, consisting of 2,990 diverse-sourced images, each equipped with a human-asked question focusing on its low-level attributes. We then measure the correctness of MLLMs on answering these questions. b) To examine the description ability of MLLMs on low-level information, we propose the LLDescribe dataset consisting of long expert-labelled golden low-level text descriptions on 499 images, and a GPT-involved comparison pipeline between outputs of MLLMs and the golden descriptions. c) Besides these two tasks, we further measure their visual quality assessment ability to align with human opinion scores. Specifically, we design a softmax-based strategy that enables MLLMs to predict quantifiable quality scores, and evaluate them on various existing image quality assessment (IQA) datasets. Our evaluation across the three abilities confirms that MLLMs possess preliminary low-level visual skills. However, these skills are still unstable and relatively imprecise, indicating the need for specific enhancements on MLLMs towards these abilities. We hope that our benchmark can encourage the research community to delve deeper to discover and enhance these untapped potentials of MLLMs. Project Page: https://q-future.github.io/Q-Bench.
翻译:多模态大语言模型(MLLMs)的快速发展推动了计算机视觉从专用模型向通用基础模型的范式转变。然而,当前在评估MLLMs底层视觉感知与理解能力方面仍存在显著不足。为填补这一空白,我们提出Q-Bench——一个系统性评估MLLMs在三个领域潜在能力的综合基准测试:底层视觉感知、底层视觉描述和整体视觉质量评估。a)为评估底层视觉感知能力,我们构建了LLVisionQA数据集,包含2,990张多源图像,每张图像配有聚焦其底层属性的人工提问,并通过衡量MLLMs回答的正确性进行评测。b)为检验MLLMs对底层信息的描述能力,我们提出LLDescribe数据集,包含499张图像及其专家标注的长文本黄金底层描述,并设计基于GPT的对比流程比较MLLMs输出与黄金描述。c)除上述两项任务外,我们进一步评估MLLMs的视觉质量感知能力以匹配人类主观评分。具体而言,我们设计基于softmax的策略使MLLMs能预测可量化的质量分数,并在多个现有图像质量评估(IQA)数据集上进行评测。对这三项能力的评估表明,MLLMs已具备初步的底层视觉技能,但这些技能仍不稳定且精度不足,亟需针对性增强。我们期望该基准测试能推动学界深入探索并提升MLLMs这些尚未被充分开发的潜力。项目页面:https://q-future.github.io/Q-Bench。