With the proliferation of hate speech on social networks under different formats, such as abusive language, cyberbullying, and violence, etc., people have experienced a significant increase in violence, putting them in uncomfortable situations and threats. Plenty of efforts have been dedicated in the last few years to overcome this phenomenon to detect hate speech in different structured languages like English, French, Arabic, and others. However, a reduced number of works deal with Arabic dialects like Tunisian, Egyptian, and Gulf, mainly the Algerian ones. To fill in the gap, we propose in this work a complete approach for detecting hate speech on online Algerian messages. Many deep learning architectures have been evaluated on the corpus we created from some Algerian social networks (Facebook, YouTube, and Twitter). This corpus contains more than 13.5K documents in Algerian dialect written in Arabic, labeled as hateful or non-hateful. Promising results are obtained, which show the efficiency of our approach.
翻译:随着仇恨言论以不同形式(如辱骂性语言、网络霸凌、暴力等)在社交网络上的泛滥,人们遭遇的暴力行为显著增加,使其陷入不适处境与威胁之中。近年来,学界为应对这一现象投入了大量努力,旨在检测不同结构语言(如英语、法语、阿拉伯语等)中的仇恨言论。然而,针对阿拉伯方言(如突尼斯方言、埃及方言、海湾方言,尤其是阿尔及利亚方言)的研究工作相对较少。为填补这一空白,本文提出了一种完整的在线阿尔及利亚语消息仇恨言论检测方法。我们在自建的阿尔及利亚社交网络(脸书、YouTube和推特)语料库上评估了多种深度学习架构。该语料库包含超过1.35万篇以阿拉伯文字书写的阿尔及利亚方言文档,标注为"仇恨性"或"非仇恨性"。实验取得了具有说服力的结果,充分证明了所提方法的有效性。