We introduce MMMU: a new benchmark designed to evaluate multimodal models on massive multi-discipline tasks demanding college-level subject knowledge and deliberate reasoning. MMMU includes 11.5K meticulously collected multimodal questions from college exams, quizzes, and textbooks, covering six core disciplines: Art & Design, Business, Science, Health & Medicine, Humanities & Social Science, and Tech & Engineering. These questions span 30 subjects and 183 subfields, comprising 30 highly heterogeneous image types, such as charts, diagrams, maps, tables, music sheets, and chemical structures. Unlike existing benchmarks, MMMU focuses on advanced perception and reasoning with domain-specific knowledge, challenging models to perform tasks akin to those faced by experts. The evaluation of 14 open-source LMMs as well as the proprietary GPT-4V(ision) and Gemini highlights the substantial challenges posed by MMMU. Even the advanced GPT-4V and Gemini Ultra only achieve accuracies of 56% and 59% respectively, indicating significant room for improvement. We believe MMMU will stimulate the community to build next-generation multimodal foundation models towards expert artificial general intelligence.
翻译:我们提出了MMMU:一个旨在评估多模态模型在需要大学水平学科知识与深思熟虑推理的大规模多学科任务上表现的新基准。MMMU包含来自大学考试、测验和教科书的11.5K个精心收集的多模态问题,涵盖六大核心学科:艺术与设计、商业、科学、健康与医学、人文与社会科学以及技术与工程。这些问题横跨30个学科和183个子领域,包含30种高度异构的图像类型,如图表、示意图、地图、表格、乐谱和化学结构。与现有基准不同,MMMU专注于结合领域特定知识的高级感知与推理,挑战模型执行类似专家面临的任务。对14个开源LMM以及专有GPT-4V(视觉版)和Gemini的评估突显了MMMU所带来的重大挑战。即使是先进的GPT-4V和Gemini Ultra,准确率也分别仅有56%和59%,表明仍有显著的改进空间。我们相信MMMU将激励社区构建面向专家级通用人工智能的下一代多模态基础模型。