As LLM benchmarks saturate, the evaluation community has pursued two strategies to increase difficulty: escalating knowledge demands (GPQA, HLE) or removing knowledge entirely in favor of abstract reasoning (ARC-AGI). The first conflates memorization with capability; the second divorces reasoning from the practical contexts in which it matters. We take a different approach. The Grounded Integration Measure (GIM) is a benchmark of 820 original problems (615 public, 205 private) where difficulty comes from integration; individual problems require coordinating multiple cognitive operations (constraint satisfaction, state tracking, epistemic vigilance, audience calibration) over broadly accessible knowledge, so that reasoning stays grounded in realistic tasks without being gated on specialized expertise. Each problem is an original expert-authored composition, majority with rubric-decomposed scoring (median 6 independently judged criteria). A balanced public--private split provides built-in contamination diagnostic. We calibrate a continuous response 2-parameter logistic (2PL) IRT model over >200k prompt-response pairs across 28 models, producing robust ability estimates that correctly order test-configurations even when raw accuracy is distorted by errors or missing data, addressing a common challenge in benchmark reporting. Using this framework, we present a comprehensive leaderboard spanning 22 models and 47 test-configurations (unique model, thinking-level pairs), and conduct what is to our knowledge the most extensive published study of how test-time compute trades off against model capability on a fixed benchmark: 11 models swept across 35 test-configurations. We observe that within-family configuration choices, such as thinking budget and quantization, matter as much as model selection. We release the evaluation framework, calibrated IRT parameters, and all public problems.


翻译:摘要:随着大语言模型(LLM)基准测试趋于饱和,评估界已采取两种策略来提升难度:一是提高知识需求(如GPQA、HLE),二是完全摒弃知识,转而依赖抽象推理(如ARC-AGI)。前者混淆了记忆与能力,后者则使推理脱离其实践意义所在的现实情境。我们另辟蹊径。基础整合度量(Ground Integration Measure, GIM)是一个包含820道原创问题(615道公开,205道私密)的基准测试,其难度源于整合:每个问题需协调多种认知操作(约束满足、状态追踪、认知警觉、受众校准),并基于广泛可及的知识,使推理扎根于现实任务,而非受限于专业知识的门槛。每道问题均由专家原创编写,多数采用评分细则分解评分(中位数包含6个独立评判标准)。公开与私密问题的平衡划分内置了污染诊断机制。我们基于28个模型的超过20万个提示-响应对,校准了连续响应双参数逻辑斯蒂克(2PL)项目反应理论(IRT)模型,生成了稳健的能力估计值,这些估计值能在原始准确率因误差或缺失数据而失真时,正确地对测试配置进行排序,从而解决了基准报告中的常见难题。利用这一框架,我们推出了涵盖22个模型和47种测试配置(各具特色的模型与思考层级组合)的综合排行榜,并开展了据我们所知关于测试时计算与模型能力在固定基准上如何权衡的最广泛公开研究:11个模型在35种测试配置上进行了扫描。我们观察到,家族内的配置选择(如思考预算和量化)与模型选择同等重要。我们已发布评估框架、校准的IRT参数及所有公开问题。

0
下载
关闭预览

相关内容

ACM/IEEE第23届模型驱动工程语言和系统国际会议,是模型驱动软件和系统工程的首要会议系列,由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来,模型涵盖了建模的各个方面,从语言和方法到工具和应用程序。模特的参加者来自不同的背景,包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛,参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会,并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。 官网链接:http://www.modelsconference.org/
评估大语言模型在科学发现中的作用
专知会员服务
19+阅读 · 2025年12月19日
大型语言模型(LLM)赋能的知识图谱构建:综述
专知会员服务
56+阅读 · 2025年10月24日
【ICML2025】通过多智能体反思强化大语言模型推理
专知会员服务
24+阅读 · 2025年6月11日
赋能大型语言模型多领域资源挑战
专知会员服务
11+阅读 · 2025年6月10日
大型语言模型推理增强外部知识:综述
专知会员服务
39+阅读 · 2025年6月2日
MME-Survey:多模态大型语言模型评估的综合性调查
专知会员服务
43+阅读 · 2024年12月1日
「知识增强预训练语言模型」最新研究综述
专知
18+阅读 · 2022年11月18日
国家自然科学基金
10+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
VIP会员
最新内容
驱动军事决策变革的顶尖人工智能指挥系统
专知会员服务
7+阅读 · 8月11日
非对称防御中的自组织临界性:俄乌战争
专知会员服务
10+阅读 · 8月10日
《战争中的大语言模型监管》
专知会员服务
14+阅读 · 8月10日
《边缘计算关键技术分析及美军作战实践应用》
边缘计算的军事应用
专知会员服务
12+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
13+阅读 · 8月8日
相关VIP内容
评估大语言模型在科学发现中的作用
专知会员服务
19+阅读 · 2025年12月19日
大型语言模型(LLM)赋能的知识图谱构建:综述
专知会员服务
56+阅读 · 2025年10月24日
【ICML2025】通过多智能体反思强化大语言模型推理
专知会员服务
24+阅读 · 2025年6月11日
赋能大型语言模型多领域资源挑战
专知会员服务
11+阅读 · 2025年6月10日
大型语言模型推理增强外部知识:综述
专知会员服务
39+阅读 · 2025年6月2日
MME-Survey:多模态大型语言模型评估的综合性调查
专知会员服务
43+阅读 · 2024年12月1日
相关基金
国家自然科学基金
10+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员