Human evaluation plays a critical role in assessing the quality of generated text. However, the reliability and reproducibility of these evaluations depend on transparent and well-documented protocols -- details that are frequently missing in current practice. In this work, we conduct a large-scale analysis of human evaluation protocols for evaluating long-form generation tasks in *CL conference publications from 2023--2025, including a full manual review of 284 papers and LLM-assisted analysis for another 1.8k+ papers. We define a set of 20 reportable criteria related to reproducibility of human evaluation studies, and apply these criteria to systematically examine reporting norms and practices within the community. We find widespread under-reporting of important aspects of human evaluation study design, leading to ambiguity about what was measured and how, who contributed judgments, and how judgments should be interpreted. Based on these findings, we outline actionable recommendations to support more transparent and reproducible reporting in future research. Our analysis code and annotated dataset can be found at: https://github.com/larchlab/Illusions-of-the-Gold-Standard
翻译:人工评估在评估生成文本质量中扮演着关键角色。然而,这些评估的可靠性和可复现性依赖于透明且记录完善的流程——这些细节在当前实践中经常缺失。在本工作中,我们对2023-2025年期间*CL会议论文中用于评估长篇生成任务的人工评估流程进行了大规模分析,包括对284篇论文的完整人工审阅,以及基于LLM辅助分析的1800余篇论文。我们定义了20项与人工评估研究可复现性相关的可报告标准,并应用这些标准系统性地审视了学术社区的汇报规范与实践。我们发现,人工评估研究设计中重要方面的汇报普遍不足,导致对测量内容、测量方式、评判贡献者以及评判解释方式产生了模糊性。基于这些发现,我们提出了可操作的建议,以支持未来研究中更透明、更可复现的汇报。我们的分析代码与标注数据集可在以下网址获取:https://github.com/larchlab/Illusions-of-the-Gold-Standard