Although Large Language Models (LLMs) and Large Multimodal Models (LMMs) exhibit impressive skills in various domains, their ability for mathematical reasoning within visual contexts has not been formally examined. Equipping LLMs and LMMs with this capability is vital for general-purpose AI assistants and showcases promising potential in education, data analysis, and scientific discovery. To bridge this gap, we present MathVista, a benchmark designed to amalgamate challenges from diverse mathematical and visual tasks. We first taxonomize the key task types, reasoning skills, and visual contexts from the literature to guide our selection from 28 existing math-focused and visual question answering datasets. Then, we construct three new datasets, IQTest, FunctionQA, and PaperQA, to accommodate for missing types of visual contexts. The problems featured often require deep visual understanding beyond OCR or image captioning, and compositional reasoning with rich domain-specific tools, thus posing a notable challenge to existing models. We conduct a comprehensive evaluation of 11 prominent open-source and proprietary foundation models (LLMs, LLMs augmented with tools, and LMMs), and early experiments with GPT-4V. The best-performing model, Multimodal Bard, achieves only 58% of human performance (34.8% vs 60.3%), indicating ample room for further improvement. Given this significant gap, MathVista fuels future research in the development of general-purpose AI agents capable of tackling mathematically intensive and visually rich real-world tasks. Preliminary tests show that MathVista also presents challenges to GPT-4V, underscoring the benchmark's importance. The project is available at https://mathvista.github.io/.
翻译:尽管大语言模型(LLMs)和多模态大模型(LMMs)在多个领域展现出令人瞩目的技能,但它们在视觉上下文中的数学推理能力尚未得到正式检验。赋予LLMs和LMMs这一能力对于通用人工智能助手至关重要,并在教育、数据分析和科学发现中展现出巨大潜力。为填补这一空白,我们提出了MathVista——一个旨在整合多样化数学与视觉任务挑战的基准测试。我们首先从文献中对关键任务类型、推理技能和视觉上下文进行分类,以指导从28个现有数学聚焦和视觉问答数据集中进行筛选。随后,我们构建了三个新数据集——IQTest、FunctionQA和PaperQA,以弥补缺失的视觉上下文类型。所涉及的问题通常需要超越OCR或图像描述的深度视觉理解,以及借助丰富领域特定工具的组合推理,这对现有模型构成了显著挑战。我们对11个主流开源和专有基础模型(包括LLMs、工具增强的LLMs以及LMMs)进行了全面评估,并对GPT-4V进行了初步实验。表现最佳的模型Multimodal Bard仅达到人类性能的58%(34.8%对比60.3%),表明仍有巨大改进空间。鉴于这一显著差距,MathVista推动了未来开发能够处理数学密集型和视觉丰富现实任务的通用AI智能体的研究。初步测试表明,MathVista同样对GPT-4V构成挑战,突显了该基准测试的重要性。项目地址:https://mathvista.github.io/。