LLM leaderboards are widely used to compare models and guide deployment decisions. However, leaderboard rankings are shaped by evaluation priorities set by benchmark designers, rather than by the diverse goals and constraints of actual users and organizations. A single aggregate score often obscures how models behave across different prompt types and compositions. In this work, we conduct an in-depth analysis of the dataset used in the LMArena (formerly Chatbot Arena) benchmark and investigate this evaluation challenge by designing an interactive visualization interface as a design probe. Our analysis reveals that the dataset is heavily skewed toward certain topics, that model rankings vary across prompt slices, and that preference-based judgments are used in ways that blur their intended scope. Building on this analysis, we introduce a visualization interface that allows users to define their own evaluation priorities by selecting and weighting prompt slices and to explore how rankings change accordingly. A qualitative study suggests that this interactive approach improves transparency and supports more context-specific model evaluation, pointing toward alternative ways to design and use LLM leaderboards.
翻译:LLM排行榜被广泛用于比较模型并指导部署决策。然而,排行榜排名由基准测试设计者设定的评估优先级决定,而非依据实际用户和组织多样化的目标与约束。单一聚合分数往往掩盖了模型在不同提示类型和组成下的表现差异。本研究对LMArena(原Chatbot Arena)基准测试中使用的数据集进行深度分析,并设计交互式可视化界面作为设计探测工具来探究这一评估挑战。分析表明:该数据集在特定主题上存在严重偏差,模型排名随提示切片变化,且基于偏好的判断在使用方式上模糊了其预期范围。基于此分析,我们引入一个可视化界面,允许用户通过选择和加权提示切片来定义自己的评估优先级,并探索排名如何随之变化。一项定性研究表明,这种交互式方法提升了透明度,支持更符合具体情境的模型评估,为设计和应用LLM排行榜提供了替代方案。