Large language model (LLM) leaderboards rank AI models using standardized benchmarks and have become highly visible across computer science, despite known limitations in their reliability and robustness. Yet how they shape researchers' actual practice remains empirically uncharted. We address this gap through semi-structured interviews with eight researchers across four computer science subfields, analyzed using reflexive thematic analysis. We find a near-universal paradox of pragmatic skepticism: while participants expressed deep distrust of leaderboard rankings, they continued to use them as rough decision-making aids. Peer networks, not leaderboards, emerged as the primary model selection mechanism, and arena-based (human-voting) leaderboards were consistently preferred over static benchmark leaderboards. Leaderboard influence varied sharply across subfields, revealing that disciplinary culture, not individual attitudes, mediates engagement; for instance, NLP researchers faced state-of-the-art comparison pressure while HCI and Systems/Privacy researchers reported none. Across these differences, however, participants converged on cost transparency as the most demanded missing feature (seven of eight). We translate these findings into concrete design recommendations that align evaluation infrastructure with how researchers actually use it, such as task-specific score breakdowns, cost integration, and voter-demographic disclosure.
翻译:大语言模型(LLM)排行榜利用标准化基准对AI模型进行排名,尽管其可靠性和稳健性存在已知局限,但仍已在计算机科学领域广为流传。然而,这些排行榜如何影响研究人员实际实践的经验证据仍处于空白。我们通过对四个计算机科学子领域的八名研究人员进行半结构化访谈(采用反思性主题分析法)来填补这一空白。研究发现了一个近乎普遍存在的实用主义怀疑悖论:尽管参与者对排行榜排名表现出深度不信任,但他们仍将其作为粗略决策辅助工具。同行网络(而非排行榜)成为模型选择的主要机制,基于竞技场(人类投票)的排行榜始终比静态基准排行榜更受青睐。排行榜的影响力在不同子领域间差异显著,表明起调节作用的是学科文化而非个人态度:例如,自然语言处理(NLP)研究人员面临追求最先进(SOTA)的对比压力,而人机交互(HCI)与系统/隐私研究人员则未报告此类压力。然而,在这些差异中,参与者一致认为成本透明度是最亟需的缺失特性(八人中有七人)。我们将这些发现转化为具体的设计建议,使评估基础设施与研究人员实际使用方式相匹配,例如任务特定分数分解、成本集成以及投票者人口统计信息的披露。