Based on powerful Large Language Models (LLMs), recent generative Multimodal Large Language Models (MLLMs) have gained prominence as a pivotal research area, exhibiting remarkable capability for both comprehension and generation. In this work, we address the evaluation of generative comprehension in MLLMs as a preliminary step towards a comprehensive assessment of generative models, by introducing a benchmark named SEED-Bench. SEED-Bench consists of 19K multiple choice questions with accurate human annotations (x 6 larger than existing benchmarks), which spans 12 evaluation dimensions including the comprehension of both the image and video modality. We develop an advanced pipeline for generating multiple-choice questions that target specific evaluation dimensions, integrating both automatic filtering and manual verification processes. Multiple-choice questions with groundtruth options derived from human annotation enables an objective and efficient assessment of model performance, eliminating the need for human or GPT intervention during evaluation. We further evaluate the performance of 18 models across all 12 dimensions, covering both the spatial and temporal understanding. By revealing the limitations of existing MLLMs through evaluation results, we aim for SEED-Bench to provide insights for motivating future research. We will launch and consistently maintain a leaderboard to provide a platform for the community to assess and investigate model capability.
翻译:基于强大的大语言模型,近年来生成的生成式多模态大语言模型已成为一个关键研究领域,展现出卓越的理解与生成能力。在本工作中,我们通过引入名为SEED-Bench的基准测试,将评估多模态大语言模型的生成式理解能力作为迈向全面评估生成式模型的第一步。SEED-Bench包含19K道带有精确人工标注的多选题(规模比现有基准大6倍),涵盖包括图像和视频模态理解在内的12个评估维度。我们开发了一种先进的多选题生成流程,针对特定评估维度,结合自动筛选与人工验证过程。基于人工标注得到真实选项的多选题能够客观高效地评估模型性能,无需在评估过程中引入人工或GPT干预。我们进一步评估了18个模型在所有12个维度上的表现,覆盖空间与时间理解。通过评估结果揭示现有多模态大语言模型的局限性,我们期望SEED-Bench能为推动未来研究提供启示。我们将发布并持续维护一个排行榜,为社区提供评估和探究模型能力的平台。