Cross-lingual transfer has become an effective way of transferring knowledge between languages. In this paper, we explore an often overlooked aspect in this domain: the influence of the source language of a language model on language transfer performance. We consider a case where the target language and its script are not part of the pre-trained model. We conduct a series of experiments on monolingual and multilingual models that are pre-trained on different tokenization methods to determine factors that affect cross-lingual transfer to a new language with a unique script. Our findings reveal the importance of the tokenizer as a stronger factor than the shared script, language similarity, and model size.
翻译:跨语言迁移已成为语言间知识迁移的有效途径。本文探讨了该领域中一个常被忽视的方面:语言模型源语言对语言迁移性能的影响。我们考虑了目标语言及其文字不在预训练模型中的情况。通过在一系列基于不同分词方法预训练的单语和多语模型上进行实验,我们确定了影响向具有独特文字的新语言进行跨语言迁移的因素。研究结果表明,分词器的重要性超过共享文字、语言相似性和模型规模,是更强的决定性因素。