Multilingual language models are widely used to extend NLP systems to low-resource languages. However, concrete evidence for the effects of multilinguality on language modeling performance in individual languages remains scarce. Here, we pre-train over 10,000 monolingual and multilingual language models for over 250 languages, including multiple language families that are under-studied in NLP. We assess how language modeling performance in each language varies as a function of (1) monolingual dataset size, (2) added multilingual dataset size, (3) linguistic similarity of the added languages, and (4) model size (up to 45M parameters). We find that in moderation, adding multilingual data improves low-resource language modeling performance, similar to increasing low-resource dataset sizes by up to 33%. Improvements depend on the syntactic similarity of the added multilingual data, with marginal additional effects of vocabulary overlap. However, high-resource languages consistently perform worse in multilingual pre-training scenarios. As dataset sizes increase, adding multilingual data begins to hurt performance for both low-resource and high-resource languages, likely due to limited model capacity (the "curse of multilinguality"). These results suggest that massively multilingual pre-training may not be optimal for any languages involved, but that more targeted models can significantly improve performance.
翻译:多语言语言模型被广泛用于将自然语言处理(NLP)系统扩展到低资源语言。然而,关于多语言性对单一语言建模性能影响的具体证据仍然稀缺。为此,我们针对超过250种语言(包括NLP中研究不足的多个语系)预训练了超过10000个单语和多语言模型。我们评估了每种语言的语言建模性能如何随以下因素变化:(1)单语数据集规模,(2)新增多语言数据集规模,(3)新增语言的 linguistic 相似性,以及(4)模型规模(最高4500万参数)。研究发现,适度添加多语言数据能提升低资源语言建模性能,效果相当于将低资源数据集规模扩大最多33%。性能提升取决于新增多语言数据的句法相似性,而词汇重叠的边际影响较小。然而,高资源语言在多语言预训练场景中始终表现更差。随着数据集规模增长,添加多语言数据开始对低资源和高资源语言的性能产生负面影响,这可能是模型容量有限所致(即“多语言性诅咒”)。这些结果表明,大规模多语言预训练可能并非对所有相关语言最优,而更具针对性的模型能显著提升性能。