Because of its importance in studying people's thoughts on various Web 2.0 services, emotion classification (EC) is an important undertaking. Existing research, on the other hand, is mostly focused on the English language, with little work on low-resource languages. Though sentiment analysis, particularly the EC in English, has received a lot of attention in recent years, little study has been done in the context of Bangla, one of the world's most widely spoken languages. We propose a complete set of approaches for identifying and extracting emotions from Bangla texts in this research. We provide a Bangla emotion classifier for six classes (anger, disgust, fear, joy, sadness, and surprise) from Bangla words, using transformer-based models which exhibit phenomenal results in recent days, especially for high resource languages. The "Unified Bangla Multi-class Emotion Corpus (UBMEC)" is used to assess the performance of our models. UBMEC was created by combining two previously released manually labeled datasets of Bangla comments on 6-emotion classes with fresh manually tagged Bangla comments created by us. The corpus dataset and code we used in this work is publicly available.
翻译:由于其在研究人们对各种Web 2.0服务观点中的重要性,情感分类(EC)是一项重要的研究任务。然而,现有研究主要集中于英语,针对低资源语言的工作较少。尽管情感分析,特别是英语中的情感分类,近年来受到广泛关注,但在全球使用最广泛的语言之一——孟加拉语的语境中,相关研究仍十分有限。本研究提出了一套完整的孟加拉语文本情感识别与提取方法。我们基于近期在高资源语言中展现出卓越性能的Transformer模型,开发了一种面向六类情感(愤怒、厌恶、恐惧、快乐、悲伤和惊讶)的孟加拉语情感分类器。我们使用“统一孟加拉语多类情感语料库(UBMEC)”评估模型性能。该语料库由两部分融合而成:两个先前发布的基于六类情感标签的孟加拉语评论人工标注数据集,以及我们新创建的人工标注孟加拉语评论数据集。本研究所用的语料库数据集及代码均已公开。