Lack of available resources such as text corpora for low-resource languages seriously hinders research on natural language processing and computational linguistics. This paper presents AlbMoRe, a corpus of 800 sentiment annotated movie reviews in Albanian. Each text is labeled as positive or negative and can be used for sentiment analysis research. Preliminary results based on traditional machine learning classifiers trained with the AlbMoRe samples are also reported. They can serve as comparison baselines for future research experiments.
翻译:低资源语言文本语料库等可用资源的匮乏严重阻碍了自然语言处理与计算语言学领域的研究。本文介绍了AlbMoRe——一个包含800条阿尔巴尼亚语情感标注电影评论的语料库。每条文本被标注为正面或负面情感,可用于情感分析研究。此外,本文还报告了基于传统机器学习分类器在AlbMoRe样本上训练的初步实验结果,这些结果可作为未来研究实验的对比基线。