Reference-based metrics such as BLEU and BERTScore are widely used to evaluate question generation (QG). In this study, on QG benchmarks such as SQuAD and HotpotQA, we find that using human-written references cannot guarantee the effectiveness of the reference-based metrics. Most QG benchmarks have only one reference; we replicated the annotation process and collect another reference. A good metric was expected to grade a human-validated question no worse than generated questions. However, the results of reference-based metrics on our newly collected reference disproved the metrics themselves. We propose a reference-free metric consisted of multi-dimensional criteria such as naturalness, answerability, and complexity, utilizing large language models. These criteria are not constrained to the syntactic or semantic of a single reference question, and the metric does not require a diverse set of references. Experiments reveal that our metric accurately distinguishes between high-quality questions and flawed ones, and achieves state-of-the-art alignment with human judgment.
翻译:基于参考的指标(如BLEU和BERTScore)被广泛用于评估问题生成(QG)。本研究发现,在SQuAD和HotpotQA等QG基准测试中,使用人工撰写的参考问题无法保证基于参考的指标的有效性。大多数QG基准仅包含一个参考问题;我们复刻了标注流程并收集了另一个参考问题。一个优质的指标理应判定人工验证过的问题质量不低于生成的问题。然而,基于参考的指标在我们新收集的参考问题上得出的结果却否定了指标本身。我们提出了一种无参考指标,该指标由自然度、可回答性和复杂度等多维标准组成,并利用大语言模型实现。这些标准不受单一参考问题句法或语义的束缚,且该指标无需多样化的参考集合。实验表明,我们的指标能够准确区分高质量问题与有缺陷的问题,并与人类判断实现了最先进的一致性。