The widespread popularity of social media has led to an increase in hateful, abusive, and sexist language, motivating methods for the automatic detection of such phenomena. The goal of the SemEval shared task \textit{Towards Explainable Detection of Online Sexism} (EDOS 2023) is to detect sexism in English social media posts (subtask A), and to categorize such posts into four coarse-grained sexism categories (subtask B), and eleven fine-grained subcategories (subtask C). In this paper, we present our submitted systems for all three subtasks, based on a multi-task model that has been fine-tuned on a range of related tasks and datasets before being fine-tuned on the specific EDOS subtasks. We implement multi-task learning by formulating each task as binary pairwise text classification, where the dataset and label descriptions are given along with the input text. The results show clear improvements over a fine-tuned DeBERTa-V3 serving as a baseline leading to $F_1$-scores of 85.9\% in subtask A (rank 13/84), 64.8\% in subtask B (rank 19/69), and 44.9\% in subtask C (26/63).
翻译:社交媒体的广泛普及导致了仇恨、辱骂和性别歧视语言的增加,这促使人们研究自动检测此类现象的方法。SemEval共享任务《面向可解释的在线性别歧视检测》(EDOS 2023)的目标是检测英文社交媒体帖子中的性别歧视(子任务A),并将此类帖子分为四个粗粒度性别歧视类别(子任务B)和十一个细粒度子类别(子任务C)。本文中,我们提出了针对所有三个子任务的提交系统,该系统基于一个多任务模型,该模型在一系列相关任务和数据集上进行了微调,然后针对特定的EDOS子任务进一步微调。我们通过将每个任务形式化为二元成对文本分类来实现多任务学习,在输入文本的同时提供数据集和标签描述。结果显示,与作为基线的微调DeBERTa-V3相比,该模型在子任务A中取得了85.9%的F1分数(排名13/84),子任务B中为64.8%(排名19/69),子任务C中为44.9%(排名26/63),表现出显著改进。