Offensive language such as hate, abuse, and profanity (HAP) occurs in various content on the web. While previous work has mostly dealt with sentence level annotations, there have been a few recent attempts to identify offensive spans as well. We build upon this work and introduce Muted, a system to identify multilingual HAP content by displaying offensive arguments and their targets using heat maps to indicate their intensity. Muted can leverage any transformer-based HAP-classification model and its attention mechanism out-of-the-box to identify toxic spans, without further fine-tuning. In addition, we use the spaCy library to identify the specific targets and arguments for the words predicted by the attention heatmaps. We present the model's performance on identifying offensive spans and their targets in existing datasets and present new annotations on German text. Finally, we demonstrate our proposed visualization tool on multilingual inputs.
翻译:摘要:仇恨、辱骂和亵渎(HAP)等攻击性语言在网络上各类内容中普遍存在。以往研究多聚焦于句子级标注,近期虽有少数工作尝试识别攻击性片段,但尚不充分。本文在此基础之上提出Muted系统——一种通过热力图直观展示攻击性论点及其攻击目标强度的多语言HAP内容识别系统。Muted可直接利用任意基于Transformer的HAP分类模型及其注意力机制识别毒性片段,无需额外微调。此外,我们采用spaCy库定位注意力热力图预测词对应的具体攻击目标及论点。我们分别在现有数据集与新标注德语文本上评估了模型对攻击性片段及其攻击目标的识别性能,最后通过多语言输入案例展示了所提出的可视化工具。