Human evaluation plays a critical role in assessing the quality of generated text. However, the reliability and reproducibility of these evaluations depend on transparent and well-documented protocols -- details that are frequently missing in current practice. In this work, we conduct a large-scale analysis of human evaluation protocols for evaluating long-form generation tasks in *CL conference publications from 2023--2025, including a full manual review of 284 papers and LLM-assisted analysis for another 1.8k+ papers. We define a set of 20 reportable criteria related to reproducibility of human evaluation studies, and apply these criteria to systematically examine reporting norms and practices within the community. We find widespread under-reporting of important aspects of human evaluation study design, leading to ambiguity about what was measured and how, who contributed judgments, and how judgments should be interpreted. Based on these findings, we outline actionable recommendations to support more transparent and reproducible reporting in future research. Our analysis code and annotated dataset can be found at: https://github.com/larchlab/Illusions-of-the-Gold-Standard


翻译:人工评估在评估生成文本质量中扮演着关键角色。然而,这些评估的可靠性和可复现性依赖于透明且记录完善的流程——这些细节在当前实践中经常缺失。在本工作中,我们对2023-2025年期间*CL会议论文中用于评估长篇生成任务的人工评估流程进行了大规模分析,包括对284篇论文的完整人工审阅,以及基于LLM辅助分析的1800余篇论文。我们定义了20项与人工评估研究可复现性相关的可报告标准,并应用这些标准系统性地审视了学术社区的汇报规范与实践。我们发现,人工评估研究设计中重要方面的汇报普遍不足,导致对测量内容、测量方式、评判贡献者以及评判解释方式产生了模糊性。基于这些发现,我们提出了可操作的建议,以支持未来研究中更透明、更可复现的汇报。我们的分析代码与标注数据集可在以下网址获取:https://github.com/larchlab/Illusions-of-the-Gold-Standard

0
下载
关闭预览

相关内容

文本、视觉与语音生成的自动化评估方法综述
专知会员服务
20+阅读 · 2025年6月15日
《AI生成视频评估综述》
专知会员服务
28+阅读 · 2024年10月30日
《生成式人工智能和情报评估》
专知会员服务
92+阅读 · 2024年7月22日
《大型语言模型自然语言生成评估》综述
专知会员服务
72+阅读 · 2024年1月20日
专知会员服务
36+阅读 · 2021年7月19日
《人工智能安全测评白皮书》,99页pdf
专知
36+阅读 · 2022年2月26日
金融领域自然语言处理研究资源大列表
专知
13+阅读 · 2020年2月27日
文本生成公开数据集/开源工具/经典论文详细列表分享
深度学习与NLP
30+阅读 · 2019年9月22日
推荐|上交大推出Texygen:文本生成模型的基准测试平台
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
8+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
VIP会员
最新内容
《基于强化学习的自动化红队测试》
专知会员服务
3+阅读 · 7月23日
伊朗不对称防空战略的演进
专知会员服务
4+阅读 · 7月23日
对抗环境下超视距目标打击的情报支援
专知会员服务
10+阅读 · 7月22日
《无人机对海面作战影响评估》
专知会员服务
15+阅读 · 7月21日
印度精确打击与指挥架构的断层
专知会员服务
7+阅读 · 7月20日
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
8+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员