While aggregate leaderboard scores drive AI development, they contain substantial measurement noise whose sources and magnitudes remain unquantified, making it unclear when rankings reflect genuine capability differences versus evaluation artifacts. We introduce a framework for measuring the latent landscape in AI benchmark ecosystems. Applying Confirmatory Factor Analysis (CFA) and Generalizability Theory to 4,000+ models from the Open LLM Leaderboard, we decompose sources of ranking variance and establish: (1) structures assumed in current reporting practice underestimate the strength of relationships between benchmarks; (2) evidence of local dependence among leaderboard items, undermining uses of benchmarks as measurement instruments under current scoring systems; (3) contributor metadata explains more rank-relevant variance ($\approx9\%$) than architecture or deployment categories in this context; (4) a manifest-score "scaling law" slope has low reliability ($R_β=0.53$); by contrast, the latent general-factor size slope is highly stable across ecosystem controls ($R_g=0.97$). We are able to provide unique insights into benchmark dynamics, such as which benchmarks are a function of LLM size and which can be oppositely impacted by post-training practices. We provide actionable diagnostics to determine how benchmark rankings can be trusted and how benchmark design can be improved.
翻译:尽管综合排行榜得分驱动着人工智能的发展,但这些得分包含大量未量化的测量噪声,其来源和程度尚不明确,因此难以判断排名反映的是真实能力差异还是评估假象。我们提出了一种用于测量AI基准测试生态系统中潜在景观的框架。通过对来自Open LLM排行榜的4000多个模型应用验证性因子分析和概括性理论,我们分解了排名差异的来源,并确立了以下结论:(1) 当前报告实践中假定的结构低估了基准测试之间的关系强度;(2) 存在排行榜项目间局部依赖的证据,这削弱了在当前评分系统下将基准测试用作测量工具的功效;(3) 在此背景下,贡献者元数据解释的排名相关方差(约9%)多于架构或部署类别;(4) 显性得分的“规模定律”斜率可靠性较低(Rβ=0.53);相比之下,潜在通用因子规模斜率在生态系统控制下高度稳定(Rg=0.97)。我们能够提供对基准测试动态的独特见解,例如哪些基准测试是LLM规模化的函数,哪些可能受到后训练实践的相反影响。我们提供了可操作的诊断方法,以确定基准测试排名在何种程度上可以被信任,以及如何改进基准测试设计。