Detecting transphobia, homophobia, and various other forms of hate speech is difficult. Signals can vary depending on factors such as language, culture, geographical region, and the particular online platform. Here, we present a joint multilingual (M-L) and language-specific (L-S) approach to homophobia and transphobic hate speech detection (HSD). M-L models are needed to catch words, phrases, and concepts that are less common or missing in a particular language and subsequently overlooked by L-S models. Nonetheless, L-S models are better situated to understand the cultural and linguistic context of the users who typically write in a particular language. Here we construct a simple and successful way to merge the M-L and L-S approaches through simple weight interpolation in such a way that is interpretable and data-driven. We demonstrate our system on task A of the 'Shared Task on Homophobia/Transphobia Detection in social media comments' dataset for homophobia and transphobic HSD. Our system achieves the best results in three of five languages and achieves a 0.997 macro average F1-score on Malayalam texts.
翻译:检测恐跨性别、恐同以及各种其他形式的仇恨言论是困难的。信号会因语言、文化、地理区域以及特定在线平台等因素而变化。在此,我们提出一种结合多语言(M-L)和特定语言(L-S)的方法用于恐同和恐跨性别仇恨言论检测。多语言模型需要捕捉那些在特定语言中不太常见或缺失,从而容易被特定语言模型忽视的词语、短语和概念。尽管如此,特定语言模型更擅长理解通常使用某种语言的用户所处的文化和语言背景。本文构建了一种简单有效的方式,通过可解释且数据驱动的权重插值方法,将多语言与特定语言方法融合。我们在“社交媒体评论中恐同/恐跨性别检测共享任务”A任务数据集上展示了我们的系统,用于恐同和恐跨性别仇恨言论检测。我们的系统在五种语言中的三种上取得了最佳结果,并在马拉雅拉姆语文本上达到了0.997的宏平均F1分数。