This paper presents KazSAnDRA, a dataset developed for Kazakh sentiment analysis that is the first and largest publicly available dataset of its kind. KazSAnDRA comprises an extensive collection of 180,064 reviews obtained from various sources and includes numerical ratings ranging from 1 to 5, providing a quantitative representation of customer attitudes. The study also pursued the automation of Kazakh sentiment classification through the development and evaluation of four machine learning models trained for both polarity classification and score classification. Experimental analysis included evaluation of the results considering both balanced and imbalanced scenarios. The most successful model attained an F1-score of 0.81 for polarity classification and 0.39 for score classification on the test sets. The dataset and fine-tuned models are open access and available for download under the Creative Commons Attribution 4.0 International License (CC BY 4.0) through our GitHub repository.
翻译:本文介绍了为哈萨克语情感分析开发的KazSAnDRA数据集,这是同类数据集中首个且规模最大的公开可用数据集。KazSAnDRA包含从多种来源收集的180,064条评论,并附有1至5分的数值评分,为顾客态度提供了定量表示。本研究还通过开发并评估四种机器学习模型(分别用于极性分类和分数分类)来推动哈萨克语情感分类的自动化。实验分析涵盖了平衡与非平衡场景下的结果评估。在测试集上,最优模型在极性分类任务中取得了0.81的F1分数,在分数分类任务中取得了0.39的F1分数。该数据集及微调模型已开放获取,可通过我们的GitHub仓库在知识共享署名4.0国际许可协议(CC BY 4.0)下下载使用。