Large language models (LLMs) have received a lot of attention in natural language processing (NLP) research because of their exceptional performance in understanding and generating human languages. However, low-resource languages are left behind due to the unavailability of resources. In this work, we focus on enhancing the LLaMA-2-Amharic model by integrating task-specific and generative datasets to improve language model performance for Amharic. We compile an Amharic instruction fine-tuning dataset and fine-tuned LLaMA-2-Amharic model. The fine-tuned model shows promising results in different NLP tasks. We open-source our dataset creation pipeline, instruction datasets, trained models, and evaluation outputs to promote language-specific studies on these models.
翻译:大型语言模型因其在理解和生成人类语言方面的卓越性能,在自然语言处理研究中备受关注。然而,由于资源匮乏,低资源语言的发展相对滞后。本研究聚焦于通过整合任务特定数据集与生成式数据集来优化LLaMA-2-Amharic模型,从而提升阿姆哈拉语的语言模型性能。我们构建了阿姆哈拉语指令微调数据集,并对LLaMA-2-Amharic模型进行了微调。微调后的模型在多项自然语言处理任务中展现出显著成效。为促进针对这些模型的特定语言研究,我们公开了数据集构建流程、指令数据集、训练模型及评估输出。