The focus of language model evaluation has transitioned towards reasoning and knowledge-intensive tasks, driven by advancements in pretraining large models. While state-of-the-art models are partially trained on large Arabic texts, evaluating their performance in Arabic remains challenging due to the limited availability of relevant datasets. To bridge this gap, we present ArabicMMLU, the first multi-task language understanding benchmark for Arabic language, sourced from school exams across diverse educational levels in different countries spanning North Africa, the Levant, and the Gulf regions. Our data comprises 40 tasks and 14,575 multiple-choice questions in Modern Standard Arabic (MSA), and is carefully constructed by collaborating with native speakers in the region. Our comprehensive evaluations of 35 models reveal substantial room for improvement, particularly among the best open-source models. Notably, BLOOMZ, mT0, LLama2, and Falcon struggle to achieve a score of 50%, while even the top-performing Arabic-centric model only achieves a score of 62.3%.
翻译:语言模型评估的重心已转向推理与知识密集型任务,这得益于大规模预训练模型的进步。尽管现有最先进模型部分使用了大型阿拉伯语文本进行训练,但由于相关数据集有限,评估其在阿拉伯语上的表现仍具挑战性。为弥合这一差距,我们提出了ArabicMMLU,这是首个面向阿拉伯语的多任务语言理解基准,其数据源自北非、黎凡特及海湾地区多个国家不同教育层次的学校考试。我们的数据集包含40项任务及14,575道现代标准阿拉伯语(MSA)多项选择题,并与当地母语者合作精心构建。对35个模型的综合评估显示,现有模型仍有显著提升空间,尤其是在最优开源模型中。值得注意的是,BLOOMZ、mT0、LLama2及Falcon模型得分均未达50%,而表现最好的阿拉伯语专用模型也仅取得62.3%的成绩。