Multilinguality is crucial for extending recent advancements in language modelling to diverse linguistic communities. To maintain high performance while representing multiple languages, multilingual models ideally align representations, allowing what is learned in one language to generalise to others. Prior research has emphasised the importance of parallel data and shared vocabulary elements as key factors for such alignment. In this study, we investigate an unintuitive novel driver of cross-lingual generalisation: language imbalance. In controlled experiments on perfectly equivalent cloned languages, we observe that the existence of a predominant language during training boosts the performance of less frequent languages and leads to stronger alignment of model representations across languages. Furthermore, we find that this trend is amplified with scale: with large enough models or long enough training, we observe that bilingual training data with a 90/10 language split yields better performance on both languages than a balanced 50/50 split. Building on these insights, we design training schemes that can improve performance in all cloned languages, even without altering the training data. As we extend our analysis to real languages, we find that infrequent languages still benefit from frequent ones, yet whether language imbalance causes cross-lingual generalisation there is not conclusive.
翻译:多语言性对于将语言建模的最新进展推广至不同语言社区至关重要。为在表达多种语言的同时保持高性能,多语言模型理想地应实现表征对齐,使得在一种语言中学到的知识能够泛化到其他语言。先前研究强调了平行数据和共享词汇元素作为此类对齐关键因素的重要性。在本研究中,我们探索了跨语言泛化的一个非直观新型驱动因素:语言不平衡。在完全等价克隆语言的受控实验中,我们观察到训练过程中主导语言的存在能够提升低频语言的性能,并使得不同语言的模型表征对齐更紧密。此外,我们发现该趋势随规模扩大而增强:当模型足够大或训练时间足够长时,采用90/10语言分割的双语训练数据在两种语言上的表现甚至优于平衡的50/50分割。基于这些发现,我们设计了无需改变训练数据即可提升所有克隆语言性能的训练方案。当我们将分析扩展至真实语言时,观察到低频语言仍能从高频语言中获益,但语言不平衡是否能在此场景下引发跨语言泛化尚未得到确定性结论。