Advanced Large Multimodal Models (LMMs) have demonstrated impressive performance in K-12 reasoning tasks, exhibiting great promise as intelligent tutors. Realizing this potential requires models to navigate real-world examinations effectively, yet most existing benchmarks fail to capture the complexity of authentic testing environments. Specifically, most datasets are static, prone to data contamination, and are often confined to restricted modalities, disciplines, and evaluation criteria. To address these issues, we introduce LiveK12Bench, a dynamic, holistic, multi-disciplinary benchmark designed to evaluate the reasoning abilities of LMMs in realistic examination scenarios. LiveK12Bench comprises 2K+ verified questions spanning Mathematics, Physics, Chemistry, and Biology, sourced from the latest real-world exam papers and designed to grow over time. Our framework features several core innovations: 1) featuring an automated pipeline that continuously ingests and parses the latest examination papers to mitigate data leakage; and 2) proposing a novel `Mock Exam' evaluation scheme, which assesses the ability to complete end-to-end exams autonomously with accurate and efficient reasoning paths. Extensive experiments on 12 LMMs reveal that advanced models suffer substantial performance degradation under exam-realistic constraints: GPT-5's score drops from 79 to 53 (out of 100) when process rigor and efficiency are jointly evaluated. Our findings expose critical vulnerabilities, such as sensitivity to complex visual layouts, highlighting the gap between idealized reasoning capabilities and true educational readiness. Both code and dataset are publicly available.


翻译:先进的大型多模态模型(LMMs)在K-12推理任务中展现了令人瞩目的性能,有望成为智能导师。要实现这一潜力,模型必须能有效应对真实世界的考试,但现有的大多数基准测试难以捕捉真实考试环境的复杂性。具体而言,多数数据集是静态的,容易受数据污染影响,且常局限于特定的模态、学科和评估标准。为解决这些问题,我们提出LiveK12Bench——一个动态、全面、多学科的基准测试,旨在评估LMMs在真实考试场景中的推理能力。LiveK12Bench包含2000余道经过验证的题目,涵盖数学、物理、化学和生物学科,题目源自最新的真实考试试卷,并计划随时间持续扩展。我们的框架包含几项核心创新:1)基于一个自动化流水线,持续获取并解析最新考试试卷,以缓解数据泄露问题;2)提出新颖的“模拟考试”评估方案,评估模型自主完成端到端考试并具备准确、高效推理路径的能力。在12个LMMs上进行的大量实验表明,先进模型在贴近真实考试的约束下性能显著下降:当同时评估过程严谨性和效率时,GPT-5的分数从79分降至53分(满分100分)。我们的发现揭示了关键脆弱性,例如对复杂视觉布局的敏感性,凸显了理想化推理能力与真实教育就绪度之间的差距。代码和数据集均已公开。

0
下载
关闭预览

相关内容

【CVPR2025教程】大规模多模态模型的评估:挑战与方法
专知会员服务
15+阅读 · 2025年6月13日
当持续学习遇上多模态大型语言模型:综述
专知会员服务
33+阅读 · 2025年3月5日
【博士论文】高效且有效的基础大型多模态模型学习
专知会员服务
41+阅读 · 2024年10月21日
《高效多模态大型语言模型》综述
专知会员服务
73+阅读 · 2024年5月20日
深度多模态表示学习综述论文,22页pdf
专知
33+阅读 · 2020年6月21日
国家自然科学基金
5+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Arxiv
25+阅读 · 2023年6月23日
VIP会员
最新内容
失去控制的指挥:人工智能时代的任务式指挥
专知会员服务
2+阅读 · 今天15:00
美国的新国家安全科技战略思考
专知会员服务
1+阅读 · 今天14:54
综述 | 面向大模型智能体的图结构个性化记忆
专知会员服务
4+阅读 · 9月10日
人工智能与未来空战管理
专知会员服务
6+阅读 · 9月9日
相关基金
国家自然科学基金
5+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员