While aggregate leaderboard scores drive AI development, they contain substantial measurement noise whose sources and magnitudes remain unquantified, making it unclear when rankings reflect genuine capability differences versus evaluation artifacts. We introduce a framework for measuring the latent landscape in AI benchmark ecosystems. Applying Confirmatory Factor Analysis (CFA) and Generalizability Theory to 4,000+ models from the Open LLM Leaderboard, we decompose sources of ranking variance and establish: (1) structures assumed in current reporting practice underestimate the strength of relationships between benchmarks; (2) evidence of local dependence among leaderboard items, undermining uses of benchmarks as measurement instruments under current scoring systems; (3) contributor metadata explains more rank-relevant variance ($\approx9\%$) than architecture or deployment categories in this context; (4) a manifest-score "scaling law" slope has low reliability ($R_β=0.53$); by contrast, the latent general-factor size slope is highly stable across ecosystem controls ($R_g=0.97$). We are able to provide unique insights into benchmark dynamics, such as which benchmarks are a function of LLM size and which can be oppositely impacted by post-training practices. We provide actionable diagnostics to determine how benchmark rankings can be trusted and how benchmark design can be improved.


翻译:尽管综合排行榜得分驱动着人工智能的发展,但这些得分包含大量未量化的测量噪声,其来源和程度尚不明确,因此难以判断排名反映的是真实能力差异还是评估假象。我们提出了一种用于测量AI基准测试生态系统中潜在景观的框架。通过对来自Open LLM排行榜的4000多个模型应用验证性因子分析和概括性理论,我们分解了排名差异的来源,并确立了以下结论:(1) 当前报告实践中假定的结构低估了基准测试之间的关系强度;(2) 存在排行榜项目间局部依赖的证据,这削弱了在当前评分系统下将基准测试用作测量工具的功效;(3) 在此背景下,贡献者元数据解释的排名相关方差(约9%)多于架构或部署类别;(4) 显性得分的“规模定律”斜率可靠性较低(Rβ=0.53);相比之下,潜在通用因子规模斜率在生态系统控制下高度稳定(Rg=0.97)。我们能够提供对基准测试动态的独特见解,例如哪些基准测试是LLM规模化的函数,哪些可能受到后训练实践的相反影响。我们提供了可操作的诊断方法,以确定基准测试排名在何种程度上可以被信任,以及如何改进基准测试设计。

0
下载
关闭预览

相关内容

AI 智能体系统:体系架构、应用场景及评估范式
AI4Research:科学研究中的人工智能综述
专知会员服务
38+阅读 · 2025年7月4日
首篇「Test-Time Scaling」全景综述,深入剖析AI深思之道
专知会员服务
15+阅读 · 2025年5月14日
智能遥感:AI 赋能遥感技术
专知会员服务
85+阅读 · 2022年5月29日
AI综述专栏 | 基于深度学习的目标检测算法综述
人工智能前沿讲习班
12+阅读 · 2018年12月7日
AI综述专栏|多模态学习研究进展综述
人工智能前沿讲习班
64+阅读 · 2018年7月13日
AI综述专栏 | 跨领域推荐系统文献综述(上)
人工智能前沿讲习班
13+阅读 · 2018年5月16日
全解:目标检测,图像分类、分割、生成……
全球人工智能
20+阅读 · 2017年9月15日
没错!卷积神经网络实现图像识别,就这么简单!
全球人工智能
20+阅读 · 2017年8月15日
国家自然科学基金
3+阅读 · 2017年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
28+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Arxiv
0+阅读 · 6月15日
VIP会员
最新内容
面向2027年及未来的海军情报改革
专知会员服务
2+阅读 · 8月5日
《无人机蜂群:释放人类-蜂群编队的潜能》
专知会员服务
4+阅读 · 8月5日
《战略战术化:一项综合性述评》
专知会员服务
2+阅读 · 8月5日
相关基金
国家自然科学基金
3+阅读 · 2017年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
28+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员