Recent advancements have seen Large Language Models (LLMs) and Large Multimodal Models (LMMs) surpassing general human capabilities in various tasks, approaching the proficiency level of human experts across multiple domains. With traditional benchmarks becoming less challenging for these models, new rigorous challenges are essential to gauge their advanced abilities. In this work, we present OlympiadBench, an Olympiad-level bilingual multimodal scientific benchmark, featuring 8,952 problems from Olympiad-level mathematics and physics competitions, including the Chinese college entrance exam. Each problem is detailed with expert-level annotations for step-by-step reasoning. Evaluating top-tier models on OlympiadBench, we implement a comprehensive assessment methodology to accurately evaluate model responses. Notably, the best-performing model, GPT-4V, attains an average score of 17.23% on OlympiadBench, with a mere 11.28% in physics, highlighting the benchmark rigor and the intricacy of physical reasoning. Our analysis orienting GPT-4V points out prevalent issues with hallucinations, knowledge omissions, and logical fallacies. We hope that our challenging benchmark can serve as a valuable resource for helping future AGI research endeavors.
翻译:近年来,大语言模型(LLMs)与大型多模态模型(LMMs)在多项任务中已超越人类常规能力,在多领域接近专家水平。随着传统基准对这类模型逐渐失去挑战性,需要设计新的严格评测标准来检验其高阶能力。本文提出奥林匹克竞赛基准(OlympiadBench)——一个奥赛级别的双语多模态科学基准,包含来自数学与物理奥林匹克竞赛(含中国高考)的8,952道题目。每道题均配备专家级分步推理标注。通过评估顶级模型在OlympiadBench上的表现,我们采用综合评估体系精确量化模型作答质量。值得注意的是,最佳模型GPT-4V在OlympiadBench上平均得分仅为17.23%,其中物理类题目更仅有11.28%,凸显了该基准的严苛性及物理推理的复杂性。针对GPT-4V的定向分析揭示了其存在幻觉、知识缺失及逻辑谬误等典型问题。我们期待这一高挑战性基准能为未来通用人工智能研究提供宝贵资源。