Recent studies have exploited advanced generative language models to generate Natural Language Explanations (NLE) for why a certain text could be hateful. We propose the Chain of Explanation (CoE) Prompting method, using the heuristic words and target group, to generate high-quality NLE for implicit hate speech. We improved the BLUE score from 44.0 to 62.3 for NLE generation by providing accurate target information. We then evaluate the quality of generated NLE using various automatic metrics and human annotations of informativeness and clarity scores.
翻译:近期研究利用先进的生成式语言模型生成自然语言解释(NLE),以阐明特定文本为何具有仇恨性。我们提出解释链(CoE)提示方法,通过启发式词与目标群体,为隐式仇恨言论生成高质量的NLE。通过提供准确的目标信息,我们将NLE生成的BLUE分数从44.0提升至62.3。随后,我们采用多种自动评估指标以及人工标注的信息量评分与清晰度评分,对生成的NLE质量进行了评估。