Multilingual large language models (MLLMs) are jointly trained on data from many different languages such that representation of individual languages can benefit from other languages' data. Impressive performance on zero-shot cross-lingual transfer shows that these models are capable of exploiting data from other languages. Yet, it remains unclear to what extent, and under which conditions, languages rely on each other's data. In this study, we use TracIn (Pruthi et al., 2020), a training data attribution (TDA) method, to retrieve the most influential training samples seen during multilingual fine-tuning for a particular test language. This allows us to analyse cross-lingual sharing mechanisms of MLLMs from a new perspective. While previous work studied cross-lingual sharing at the level of model parameters, we present the first approach to study cross-lingual sharing at the data level. We find that MLLMs rely on data from multiple languages from the early stages of fine-tuning and that this reliance gradually increases as fine-tuning progresses. We further study how different fine-tuning languages influence model performance on a given test language and find that they can both reinforce and complement the knowledge acquired from data of the test language itself.
翻译:多语言大语言模型(MLLMs)在来自多种不同语言的数据上联合训练,使得单个语言的表示能够受益于其他语言的数据。其在零样本跨语言迁移任务上表现出的卓越性能表明,这些模型能够有效利用其他语言的数据。然而,目前仍不清楚语言在多大程度上以及何种条件下会依赖彼此的数据。本研究采用TracIn(Pruthi等人,2020)——一种训练数据归因(TDA)方法,来检索多语言微调过程中对特定测试语言最具影响力的训练样本,从而从全新视角分析MLLMs的跨语言共享机制。以往研究在模型参数层面考察跨语言共享,而本研究首次提出在数据层面分析跨语言共享的方法。我们发现,MLLMs从微调早期阶段便开始依赖多种语言的数据,且这种依赖程度随微调进程逐步加深。我们进一步探究不同微调语言如何影响模型在给定测试语言上的表现,并发现这些语言既能强化、也能补充从测试语言自身数据中习得的知识。