Pretraining monolingual language models have been proven to be vital for performance in Arabic Natural Language Processing (NLP) tasks. In this paper, we conduct a comprehensive study on the role of data in Arabic Pretrained Language Models (PLMs). More precisely, we reassess the performance of a suite of state-of-the-art Arabic PLMs by retraining them on massive-scale, high-quality Arabic corpora. We have significantly improved the performance of the leading Arabic encoder-only BERT-base and encoder-decoder T5-base models on the ALUE and ORCA leaderboards, thereby reporting state-of-the-art results in their respective model categories. In addition, our analysis strongly suggests that pretraining data by far is the primary contributor to performance, surpassing other factors. Our models and source code are publicly available at https://github.com/huawei-noah/Pretrained-Language-Model/tree/master/JABER-PyTorch.
翻译:预训练单语言模型已被证明对阿拉伯语自然语言处理(NLP)任务的性能至关重要。本文系统研究了数据在阿拉伯语预训练语言模型(PLMs)中的作用。具体而言,我们通过在大规模、高质量阿拉伯语语料库上重新训练,重新评估了一系列当前最优阿拉伯语PLMs的性能。我们显著提升了领先的阿拉伯语仅编码器BERT-base模型和编码器-解码器T5-base模型在ALUE与ORCA排行榜上的表现,从而在各自模型类别中报告了最优结果。此外,我们的分析强烈表明,预训练数据是性能的首要贡献因素,其重要性远超其他因素。我们的模型和源代码已公开于https://github.com/huawei-noah/Pretrained-Language-Model/tree/master/JABER-PyTorch。