We present the first Africentric SemEval Shared task, Sentiment Analysis for African Languages (AfriSenti-SemEval) - the dataset is available at https://github.com/afrisenti-semeval/afrisent-semeval-2023. AfriSenti-SemEval is a sentiment classification challenge in 14 African languages - Amharic, Algerian Arabic, Hausa, Igbo, Kinyarwanda, Moroccan Arabic, Mozambican Portuguese, Nigerian Pidgin, Oromo, Swahili, Tigrinya, Twi, Xitsonga, and Yor\`ub\'a (Muhammad et al., 2023), using a 3-class labeled data: positive, negative, and neutral. We present three subtasks: (1) Task A: monolingual classification, which received 44 submissions; (2) Task B: multilingual classification, which received 32 submissions; and (3) Task C: zero-shot classification, which received 34 submissions. The best system for tasks A and B was achieved by NLNDE team with 71.31 and 75.06 weighted F1, respectively. UCAS-IIE-NLP achieved the best system on average for task C with 58.15 weighted F1. We describe the various approaches adopted by the top 10 systems and their approaches.
翻译:我们首次提出以非洲为中心的SemEval共享任务——非洲语言情感分析(AfriSenti-SemEval),数据集可于https://github.com/afrisenti-semeval/afrisent-semeval-2023获取。AfriSenti-SemEval是一项面向14种非洲语言的情感分类挑战任务,涵盖阿姆哈拉语、阿尔及利亚阿拉伯语、豪萨语、伊博语、卢旺达语、摩洛哥阿拉伯语、莫桑比克葡萄牙语、尼日利亚皮钦语、奥罗莫语、斯瓦希里语、提格雷尼亚语、契维语、聪加语和约鲁巴语(Muhammad等人,2023),采用三元标注数据:积极、消极和中性。我们提出三个子任务:(1)任务A:单语言分类,收到44份提交;(2)任务B:多语言分类,收到32份提交;(3)任务C:零样本分类,收到34份提交。任务A和B的最佳系统由NLNDE团队获得,加权F1值分别为71.31和75.06;UCAS-IIE-NLP团队在任务C中取得平均最佳系统成绩,加权F1值为58.15。我们描述了排名前10系统所采用的各种方法及其技术路径。