Similes play an imperative role in creative writing such as story and dialogue generation. Proper evaluation metrics are like a beacon guiding the research of simile generation (SG). However, it remains under-explored as to what criteria should be considered, how to quantify each criterion into metrics, and whether the metrics are effective for comprehensive, efficient, and reliable SG evaluation. To address the issues, we establish HAUSER, a holistic and automatic evaluation system for the SG task, which consists of five criteria from three perspectives and automatic metrics for each criterion. Through extensive experiments, we verify that our metrics are significantly more correlated with human ratings from each perspective compared with prior automatic metrics.
翻译:明喻在故事生成、对话生成等创意写作中发挥着不可或缺的作用。合理的评估指标如同灯塔,指引着明喻生成(SG)研究的发展方向。然而,当前研究对明喻生成评估应遵循哪些标准、如何将各标准量化为具体指标、以及这些指标能否实现全面、高效且可靠的SG评估等问题仍缺乏深入探讨。为解决上述问题,我们构建了HAUSER——面向明喻生成任务的全面自动评估体系。该体系从三个视角提出五项评估标准,并为每项标准设计了自动化评估指标。通过大量实验验证,相较于现有自动评估指标,我们提出的各项指标与人类评分之间的相关性均呈现显著提升。