We introduce EXAMS-V, a new challenging multi-discipline multimodal multilingual exam benchmark for evaluating vision language models. It consists of 20,932 multiple-choice questions across 20 school disciplines covering natural science, social science, and other miscellaneous studies, e.g., religion, fine arts, business, etc. EXAMS-V includes a variety of multimodal features such as text, images, tables, figures, diagrams, maps, scientific symbols, and equations. The questions come in 11 languages from 7 language families. Unlike existing benchmarks, EXAMS-V is uniquely curated by gathering school exam questions from various countries, with a variety of education systems. This distinctive approach calls for intricate reasoning across diverse languages and relies on region-specific knowledge. Solving the problems in the dataset requires advanced perception and joint reasoning over the text and the visual content of the image. Our evaluation results demonstrate that this is a challenging dataset, which is difficult even for advanced vision-text models such as GPT-4V and Gemini; this underscores the inherent complexity of the dataset and its significance as a future benchmark.
翻译:我们提出了EXAMS-V——一个面向视觉语言模型的挑战性多学科、多语言、多模态考试新基准。该基准包含来自20个学校学科的20,932道多项选择题,涵盖自然科学、社会科学及其他学科(如宗教、美术、商学等)。EXAMS-V集成了文本、图像、表格、图形、图表、地图、科学符号和方程等多种多模态特征。试题使用来自7个语系的11种语言呈现。与现有基准不同,EXAMS-V的独特之处在于收集了来自不同国家的学校考试试题,体现了多样化的教育体系。这种独特方法要求跨语言复杂推理,并依赖区域特定知识。解决数据集中问题需要高级感知能力,并对文本与图像视觉内容进行联合推理。我们的评估结果表明,这是一个具有挑战性的数据集,即使对GPT-4V和Gemini等先进视觉-文本模型也难以攻克,这突显了该数据集的内在复杂性及其作为未来基准的重要意义。