Africa is home to over 2000 languages from over six language families and has the highest linguistic diversity among all continents. This includes 75 languages with at least one million speakers each. Yet, there is little NLP research conducted on African languages. Crucial in enabling such research is the availability of high-quality annotated datasets. In this paper, we introduce AfriSenti, which consists of 14 sentiment datasets of 110,000+ tweets in 14 African languages (Amharic, Algerian Arabic, Hausa, Igbo, Kinyarwanda, Moroccan Arabic, Mozambican Portuguese, Nigerian Pidgin, Oromo, Swahili, Tigrinya, Twi, Xitsonga, and Yor\`ub\'a) from four language families annotated by native speakers. The data is used in SemEval 2023 Task 12, the first Afro-centric SemEval shared task. We describe the data collection methodology, annotation process, and related challenges when curating each of the datasets. We conduct experiments with different sentiment classification baselines and discuss their usefulness. We hope AfriSenti enables new work on under-represented languages. The dataset is available at https://github.com/afrisenti-semeval/afrisent-semeval-2023 and can also be loaded as a huggingface datasets (https://huggingface.co/datasets/shmuhammad/AfriSenti).
翻译:非洲拥有超过2000种语言,分属六大语系,其语言多样性居所有大陆之首。其中包含75种各拥有至少一百万使用者的语言。然而,针对非洲语言的 NLP 研究仍十分匮乏。推动这类研究的关键在于获取高质量的人工标注数据集。本文介绍了 AfriSenti,该数据集包含来自四个语系的14种非洲语言(阿姆哈拉语、阿尔及利亚阿拉伯语、豪萨语、伊博语、卢旺达语、摩洛哥阿拉伯语、莫桑比克葡萄牙语、尼日利亚皮钦语、奥罗莫语、斯瓦希里语、提格雷尼亚语、契维语、齐聪加语和约鲁巴语)的超过11万条推文所组成的14个情感数据集,并由母语者完成标注。该数据用于 SemEval 2023 的第12项任务,这是首个以非洲语言为中心的 SemEval 共享任务。我们描述了每个数据集在构建过程中的数据收集方法、标注流程及相关挑战。我们利用不同的情感分类基线进行了实验,并探讨了其有效性。希望 AfriSenti 能推动针对代表性不足语言的新研究工作。该数据集可通过 https://github.com/afrisenti-semeval/afrisent-semeval-2023 获取,也可作为 Hugging Face 数据集(https://huggingface.co/datasets/shmuhammad/AfriSenti)加载。