As the capabilities of language models continue to advance, it is conceivable that "one-size-fits-all" model will remain as the main paradigm. For instance, given the vast number of languages worldwide, many of which are low-resource, the prevalent practice is to pretrain a single model on multiple languages. In this paper, we add to the growing body of evidence that challenges this practice, demonstrating that monolingual pretraining on the target language significantly improves models already extensively trained on diverse corpora. More specifically, we further pretrain GPT-J and LLaMA models on Portuguese texts using 3% or less of their original pretraining budget. Few-shot evaluations on Poeta, a suite of 14 Portuguese datasets, reveal that our models outperform English-centric and multilingual counterparts by a significant margin. Our best model, Sabi\'a-65B, performs on par with GPT-3.5-turbo. By evaluating on datasets originally conceived in the target language as well as translated ones, we study the contributions of language-specific pretraining in terms of 1) capturing linguistic nuances and structures inherent to the target language, and 2) enriching the model's knowledge about a domain or culture. Our results indicate that the majority of the benefits stem from the domain-specific knowledge acquired through monolingual pretraining.
翻译:随着语言模型能力的持续提升,'通用型'模型可能仍将作为主要范式存在。例如,鉴于全球语言数量庞大且多数为低资源语言,当前主流做法是在多种语言上预训练单一模型。本文通过研究验证了挑战该实践的累积证据,证明在以多种语料库进行充分预训练的基础上,针对目标语言进行单语预训练可显著提升模型性能。具体而言,我们仅使用原始预训练预算的3%或更少资源,在葡萄牙语文本上进一步训练了GPT-J和LLaMA模型。基于涵盖14个葡萄牙语数据集的Poeta套件的少样本评估表明,我们的模型以显著优势超越了以英语为中心和多语言对照模型。其中最佳模型Sabiá-65B的性能与GPT-3.5-turbo相当。通过评估目标语言原生数据集和翻译数据集,我们分析了语言特定预训练的两方面贡献:1)捕获目标语言固有的语言细微特征与结构;2)丰富模型对特定领域或文化的知识储备。研究结果表明,大部分性能提升源自单语预训练带来的领域特定知识。