This paper presents a high-quality dataset for evaluating the quality of Bangla word embeddings, which is a fundamental task in the field of Natural Language Processing (NLP). Despite being the 7th most-spoken language in the world, Bangla is a low-resource language and popular NLP models fail to perform well. Developing a reliable evaluation test set for Bangla word embeddings are crucial for benchmarking and guiding future research. We provide a Mikolov-style word analogy evaluation set specifically for Bangla, with a sample size of 16678, as well as a translated and curated version of the Mikolov dataset, which contains 10594 samples for cross-lingual research. Our experiments with different state-of-the-art embedding models reveal that Bangla has its own unique characteristics, and current embeddings for Bangla still struggle to achieve high accuracy on both datasets. We suggest that future research should focus on training models with larger datasets and considering the unique morphological characteristics of Bangla. This study represents the first step towards building a reliable NLP system for the Bangla language1.
翻译:本文提出了一个用于评估孟加拉语词嵌入质量的高质量数据集,这是自然语言处理(NLP)领域的一项基础任务。尽管孟加拉语是世界第七大语言,但它仍属于低资源语言,主流NLP模型在其上的表现不佳。开发可靠的孟加拉语词嵌入评估测试集对于基准测试和指导未来研究至关重要。我们提供了一个专门针对孟加拉语的Mikolov风格词汇类比评估集,样本量为16678个,同时包含一个经过翻译和整理的Mikolov数据集版本(含10594个样本),可用于跨语言研究。通过与多种先进嵌入模型的实验,我们发现孟加拉语具有其独特特征,当前孟加拉语嵌入在两类数据集上仍难以实现高精度。我们建议未来研究应聚焦于使用更大数据集训练模型,并充分考虑孟加拉语独特的形态学特征。本研究是构建可靠孟加拉语NLP系统的第一步。