LLM-as-judge evaluation has become standard practice for open-ended model assessment; however, judges exhibit systematic biases that cannot be averaged out by increasing the number of scenarios or generations. These biases are often similar in magnitude to the model differences that benchmarks are designed to detect, resulting in unreliable rankings when single-judge evaluations are used. We introduce a variance decomposition that partitions benchmark score variance into scenario, generation, judge, and residual components. Based on this analysis, CyclicJudge, a round-robin assignment of judges to scenarios, is demonstrated to be the optimal strategy for a fixed judge panel and judge-call budget: the score recovers the panel mean exactly while matching the cost of single-judge evaluation. Empirical results on MT-Bench and MindEval validate the effectiveness of CyclicJudge as predicted, across both general-purpose and domain-specific evaluation settings.
翻译:将大模型作为判官(LLM-as-judge)的评估已成为开放性模型评测的标准实践;然而,判官存在系统性偏差,这种偏差无法通过增加评测场景或生成结果的样本量加以平均消除。此类偏差的量级往往与基准测试旨在检测的模型差异相当,导致单一判官评估结果产生不可靠的排名。我们引入一种方差分解方法,将基准测试分数方差划分为场景、生成结果、判官及残差分量。基于该分析,我们证明了循环判官法(CyclicJudge)——即对判官与场景进行循环配对分配——在固定判官面板及判官调用预算条件下为最优策略:其评分结果恰好能恢复判官面板均值,且成本与单一判官评估持平。在MT-Bench与MindEval上的实证结果验证了CyclicJudge在通用与领域特定评估场景中的预期有效性。