The proliferation of online toxic speech is a pertinent problem posing threats to demographic groups. While explicit toxic speech contains offensive lexical signals, implicit one consists of coded or indirect language. Therefore, it is crucial for models not only to detect implicit toxic speech but also to explain its toxicity. This draws a unique need for unified frameworks that can effectively detect and explain implicit toxic speech. Prior works mainly formulated the task of toxic speech detection and explanation as a text generation problem. Nonetheless, models trained using this strategy can be prone to suffer from the consequent error propagation problem. Moreover, our experiments reveal that the detection results of such models are much lower than those that focus only on the detection task. To bridge these gaps, we introduce ToXCL, a unified framework for the detection and explanation of implicit toxic speech. Our model consists of three modules: a (i) Target Group Generator to generate the targeted demographic group(s) of a given post; an (ii) Encoder-Decoder Model in which the encoder focuses on detecting implicit toxic speech and is boosted by a (iii) Teacher Classifier via knowledge distillation, and the decoder generates the necessary explanation. ToXCL achieves new state-of-the-art effectiveness, and outperforms baselines significantly.
翻译:在线有毒言论的泛滥是一个亟待解决的问题,对不同人口群体构成威胁。显性有毒言论包含具有攻击性的词汇信号,而隐性有毒言论则采用隐晦或间接的语言表达。因此,模型不仅需要检测隐性有毒言论,还需解释其毒性成因,这使得对能够有效检测并解释隐性有毒言论的统一框架需求尤为迫切。先前研究主要将有毒言论检测与解释任务建模为文本生成问题,但此类训练策略可能导致模型出现后续错误传播问题。此外,我们的实验表明,采用该策略的模型在检测结果上远低于仅专注于检测任务的模型。为弥合这些差距,我们提出了ToXCL——一个用于隐性有毒言论检测与解释的统一框架。该模型包含三个模块:(i) 目标群体生成器,用于生成给定帖子的目标人口群体;(ii) 编码器-解码器模型,其中编码器专注于检测隐性有毒言论,并通过(iii)教师分类器以知识蒸馏方式增强其性能,解码器则生成必要的解释。ToXCL实现了新的最先进效果,显著优于现有基线模型。