This study focuses on the generation of Persian named entity datasets through the application of machine translation on English datasets. The generated datasets were evaluated by experimenting with one monolingual and one multilingual transformer model. Notably, the CoNLL 2003 dataset has achieved the highest F1 score of 85.11%. In contrast, the WNUT 2017 dataset yielded the lowest F1 score of 40.02%. The results of this study highlight the potential of machine translation in creating high-quality named entity recognition datasets for low-resource languages like Persian. The study compares the performance of these generated datasets with English named entity recognition systems and provides insights into the effectiveness of machine translation for this task. Additionally, this approach could be used to augment data in low-resource language or create noisy data to make named entity systems more robust and improve them.
翻译:本研究聚焦于通过对英语数据集应用机器翻译来生成波斯语命名实体数据集。通过使用一个单语言和一个多语言Transformer模型进行实验,对所生成的数据集进行了评估。值得注意的是,CoNLL 2003数据集取得了最高F1分数85.11%,而WNUT 2017数据集则获得了最低F1分数40.02%。研究结果凸显了机器翻译在为波斯语等低资源语言创建高质量命名实体识别数据集方面的潜力。本研究将这些生成数据集的表现与英语命名实体识别系统进行了比较,并深入分析了机器翻译在此任务中的有效性。此外,该方法还可用于扩充低资源语言的数据,或生成含噪声数据,以增强命名实体系统的鲁棒性并改进其性能。