In our increasingly interconnected digital world, social media platforms have emerged as powerful channels for the dissemination of hate speech and offensive content. This work delves into the domain of hate speech detection, placing specific emphasis on three low-resource Indian languages: Bengali, Assamese, and Gujarati. The challenge is framed as a text classification task, aimed at discerning whether a tweet contains offensive or non-offensive content. Leveraging the HASOC 2023 datasets, we fine-tuned pre-trained BERT and SBERT models to evaluate their effectiveness in identifying hate speech. Our findings underscore the superiority of monolingual sentence-BERT models, particularly in the Bengali language, where we achieved the highest ranking. However, the performance in Assamese and Gujarati languages signifies ongoing opportunities for enhancement. Our goal is to foster inclusive online spaces by countering hate speech proliferation.
翻译:在这个日益互联的数字世界中,社交媒体平台已成为传播仇恨言论和攻击性内容的强大渠道。本研究深入探讨仇恨言论检测领域,特别关注三种低资源印度语言:孟加拉语、阿萨姆语和古吉拉特语。该任务被界定为文本分类问题,旨在判断推文是否包含攻击性或非攻击性内容。利用HASOC 2023数据集,我们微调了预训练的BERT和SBERT模型,以评估其在识别仇恨言论中的有效性。我们的研究结果凸显了单语句子BERT模型的优越性,尤其是在孟加拉语中,我们取得了最高排名。然而,在阿萨姆语和古吉拉特语中的表现表明仍有改进空间。我们的目标是通过遏制仇恨言论扩散来促进包容性的在线空间。