Multiple choice questions (MCQs) are an efficient and common way to assess reading comprehension (RC). Every MCQ needs a set of distractor answers that are incorrect, but plausible enough to test student knowledge. Distractor generation (DG) models have been proposed, and their performance is typically evaluated using machine translation (MT) metrics. However, MT metrics often misjudge the suitability of generated distractors. We propose DISTO: the first learned evaluation metric for generated distractors. We validate DISTO by showing its scores correlate highly with human ratings of distractor quality. At the same time, DISTO ranks the performance of state-of-the-art DG models very differently from MT-based metrics, showing that MT metrics should not be used for distractor evaluation.
翻译:多项选择题(MCQs)是评估阅读理解(RC)的一种高效且常见的方式。每道MCQ都需要一组作为错误答案的干扰项,这些干扰项需具备充分合理性以检验学生知识掌握程度。针对干扰项生成(DG)模型已有相关研究,其性能通常采用机器翻译(MT)指标进行评估。然而,MT指标常对生成干扰项的适配性产生误判。我们提出DISTO:首个针对生成干扰项的学习型评估指标。通过展示其评分与干扰项质量的人工评分高度相关,我们验证了DISTO的有效性。同时,DISTO对当前最优DG模型性能的排序结果与基于MT的指标存在显著差异,这表明不应将MT指标用于干扰项评估。