Recently, there has been a growing trend of utilizing Large Language Model (LLM) to evaluate the quality of other LLMs. Many studies have employed proprietary close-source models, especially GPT4, as the evaluator. Alternatively, other works have fine-tuned judge models based on open-source LLMs as the evaluator. In this study, we conduct an empirical study of different judge models on their evaluation capability. Our findings indicate that although the fine-tuned judge models achieve high accuracy on in-domain test sets, even surpassing GPT4, they are inherently task-specific classifiers, and their generalizability and fairness severely underperform GPT4.
翻译:近年来,利用大语言模型评估其他大语言模型质量的研究趋势日益增长。许多研究采用专有闭源模型(尤其是GPT4)作为评估器,另一些研究则基于开源大语言模型微调评判模型作为评估器。本研究对不同评判模型的评估能力进行了实证分析。结果表明:尽管微调评判模型在领域内测试集上能达到较高准确率,甚至超越GPT4,但其本质上是任务特定分类器,其泛化能力和公平性显著弱于GPT4。