In this study, we analyze NLG automatic metrics based on whether human evaluation aspect is used as context or objective to compute the metrics: (i) Task-agnostic and (ii) Human-aligned. Task-agnostic metrics, such as Perplexity, BLEU, BERTScore, are cost-effective and highly adaptable to diverse NLG tasks, yet they have a weak correlation with human. Human-aligned metrics (CTC, CtrlEval, UniEval) improves correlation level by incorporating desirable human-like qualities as training objective. However, their effectiveness at discerning system-level performance and quality of system outputs remain unclear. We present metric preference checklist as a framework to assess the discriminative power of automatic metrics in three NLG tasks: Text Summarization, Dialogue Response Generation, and Controlled Generation. We show that multi-aspect human-aligned metric (UniEval) is not necessarily dominant over single-aspect human-aligned metrics (CTC, CtrlEval) and task-agnostic metrics (BLEU, BERTScore), particularly when a disagreement between human evaluation aspects is present. We also show particular use cases in which automatic metrics provide a better guidance than human on discriminating system-level performance. Our proposed framework provides access: (i) for verifying whether automatic metrics are faithful to human preference, regardless their correlation level to human; and (ii) for scrutinizing the strengths and limitations of NLG systems, which are often obscured by a standard averaging method of evaluation scores.
翻译:本研究基于人类评估方面是否被用作计算指标的上下文或目标,对NLG自动指标进行了分析:(i) 任务无关指标 和 (ii) 人类对齐指标。任务无关指标(如Perplexity、BLEU、BERTScore)成本低廉且高度适应各种NLG任务,但与人类的相关性较弱。人类对齐指标(CTC、CtrlEval、UniEval)通过将理想的人类特质作为训练目标来提高相关性水平。然而,它们在辨别系统级性能和系统输出质量方面的有效性仍不明确。我们提出了指标偏好清单作为评估三种NLG任务中自动指标区分能力的框架:文本摘要、对话响应生成和受控生成。我们表明,多维度人类对齐指标(UniEval)不一定优于单维度人类对齐指标(CTC、CtrlEval)和任务无关指标(BLEU、BERTScore),特别是在人类评估方面存在分歧时。我们还展示了自动指标在辨别系统级性能方面比人类提供更好指导的特定用例。我们提出的框架提供了以下途径:(i) 验证自动指标是否忠实于人类偏好,无论其与人类的相关性水平如何;(ii) 审视NLG系统的优势与局限,这些优势与局限常被评估得分的标准平均方法所掩盖。