Instruction-tuned Large Language Models (LLMs) have recently showcased remarkable advancements in their ability to generate fitting responses to natural language instructions. However, many current works rely on manual evaluation to judge the quality of generated responses. Since such manual evaluation is time-consuming, it does not easily scale to the evaluation of multiple models and model variants. In this short paper, we propose a straightforward but remarkably effective evaluation metric called SemScore, in which we directly compare model outputs to gold target responses using semantic textual similarity (STS). We conduct a comparative evaluation of the model outputs of 12 prominent instruction-tuned LLMs using 8 widely-used evaluation metrics for text generation. We find that our proposed SemScore metric outperforms all other, in many cases more complex, evaluation metrics in terms of correlation to human evaluation. These findings indicate the utility of our proposed metric for the evaluation of instruction-tuned LLMs.
翻译:指令调优的大语言模型近期在生成符合自然语言指令的恰当回复方面展现了显著进步。然而,许多现有工作仍依赖人工评估来判断生成回复的质量。由于这种人工评估耗时费力,难以扩展到多个模型及变体的评估。本文提出一种简单但极为有效的评估指标SemScore,该指标通过语义文本相似度直接比较模型输出与黄金标准回复。我们利用8种广泛使用的文本生成评估指标,对12个主流指令调优大语言模型的输出进行了比较评估。研究发现,所提出的SemScore指标在人机评估一致性上优于所有其他(通常更复杂)的评估指标。这些结果证明了该指标在指令调优大语言模型评估中的实用性。