Human evaluation is the foundation upon which the evaluation of both summarization systems and automatic metrics rests. However, existing human evaluation studies for summarization either exhibit a low inter-annotator agreement or have insufficient scale, and an in-depth analysis of human evaluation is lacking. Therefore, we address the shortcomings of existing summarization evaluation along the following axes: (1) We propose a modified summarization salience protocol, Atomic Content Units (ACUs), which is based on fine-grained semantic units and allows for a high inter-annotator agreement. (2) We curate the Robust Summarization Evaluation (RoSE) benchmark, a large human evaluation dataset consisting of 22,000 summary-level annotations over 28 top-performing systems on three datasets. (3) We conduct a comparative study of four human evaluation protocols, underscoring potential confounding factors in evaluation setups. (4) We evaluate 50 automatic metrics and their variants using the collected human annotations across evaluation protocols and demonstrate how our benchmark leads to more statistically stable and significant results. The metrics we benchmarked include recent methods based on large language models (LLMs), GPTScore and G-Eval. Furthermore, our findings have important implications for evaluating LLMs, as we show that LLMs adjusted by human feedback (e.g., GPT-3.5) may overfit unconstrained human evaluation, which is affected by the annotators' prior, input-agnostic preferences, calling for more robust, targeted evaluation methods.
翻译:人工评估是摘要系统及其自动评估指标评价的基石。然而,现有摘要领域的人工评估研究存在标注者间一致性低或规模不足的问题,且缺乏深入分析。为此,我们从以下维度改进当前摘要评估的不足:(1) 提出改进的摘要显著性标注协议——原子内容单元(ACUs),该协议基于细粒度语义单元,可实现高标注者间一致性;(2) 构建稳健摘要评估基准(RoSE),包含28个顶尖系统在三个数据集上的22,000条摘要级标注数据;(3) 对四种人工评估协议开展比较研究,揭示评估设置中的潜在混淆因素;(4) 基于跨评估协议的标注数据,系统评估50种自动评估指标及其变体,证明该基准可带来统计上更稳定显著的结果。基准测试涵盖基于大语言模型(LLMs)的最新方法GPTScore与G-Eval。此外,研究发现对评估LLMs具有重要启示:经人类反馈调整的LLMs(如GPT-3.5)可能过度拟合非约束性人工评估,而此类评估易受标注者先验知识及与输入无关的偏好影响,亟需开发更具鲁棒性的定向评估方法。