Large language models (LLMs) are widely used as scalable evaluators of model responses in lieu of human annotators. However, imperfect sensitivity and specificity of the LLM judges induce bias in naive evaluation scores. We propose a simple plug-in framework that corrects this bias and enables statistically principled uncertainty quantification. Our framework constructs confidence intervals that account for uncertainty from both the test dataset and a human-labeled calibration dataset. Additionally, it uses an adaptive strategy to allocate calibration samples for tighter intervals. Importantly, we characterize parameter regimes defined by the true evaluation score and the LLM judge's sensitivity and specificity in which our LLM-based evaluation yields more reliable estimates than human-only evaluation. Moreover, we show that our framework remains unbiased under distribution shift between the test and calibration datasets, in contrast to existing approaches.
翻译:大型语言模型(LLM)被广泛用作模型响应的可扩展评估器,以替代人工标注者。然而,LLM评估器的不完美敏感性和特异性会在朴素评估分数中引入偏差。我们提出一个简单的即插即用框架,可纠正此偏差并实现具有统计学原理的不确定性量化。该框架构建的置信区间同时考虑了测试数据集和人工标注校准数据集中的不确定性。此外,它采用自适应策略分配校准样本以获得更紧凑的区间。重要的是,我们刻画了由真实评估分数和LLM评估器的敏感性与特异性定义的参数体系,在该体系中,基于LLM的评估比仅人工评估能产生更可靠的估计。最后,我们证明与现有方法不同,该框架在测试集与校准集之间存在分布偏移时仍能保持无偏性。