Without any explicit cross-lingual training data, multilingual language models can achieve cross-lingual transfer. One common way to improve this transfer is to perform realignment steps before fine-tuning, i.e., to train the model to build similar representations for pairs of words from translated sentences. But such realignment methods were found to not always improve results across languages and tasks, which raises the question of whether aligned representations are truly beneficial for cross-lingual transfer. We provide evidence that alignment is actually significantly correlated with cross-lingual transfer across languages, models and random seeds. We show that fine-tuning can have a significant impact on alignment, depending mainly on the downstream task and the model. Finally, we show that realignment can, in some instances, improve cross-lingual transfer, and we identify conditions in which realignment methods provide significant improvements. Namely, we find that realignment works better on tasks for which alignment is correlated with cross-lingual transfer when generalizing to a distant language and with smaller models, as well as when using a bilingual dictionary rather than FastAlign to extract realignment pairs. For example, for POS-tagging, between English and Arabic, realignment can bring a +15.8 accuracy improvement on distilmBERT, even outperforming XLM-R Large by 1.7. We thus advocate for further research on realignment methods for smaller multilingual models as an alternative to scaling.
翻译:在没有显式跨语言训练数据的情况下,多语言语言模型能够实现跨语言迁移。提升这种迁移的常见方法是在微调前执行重对齐步骤,即训练模型为翻译句对中的词语构建相似表示。但研究发现这类重对齐方法并非总能提升所有语言和任务的结果,这引发了对齐表示是否真正有益于跨语言迁移的疑问。我们提供证据表明,在语言、模型和随机种子维度上,对齐与跨语言迁移实际上存在显著相关性。我们展示了微调会对对齐产生显著影响,这种影响主要取决于下游任务和模型。最后,我们发现重对齐在某些情况下能够改善跨语言迁移,并识别了重对齐方法带来显著改进的条件。具体而言,当跨语言迁移与对齐存在相关性时(如表征泛化到远距离语言、使用较小模型的情况),以及使用双语词典而非FastAlign提取重对齐对时,重对齐方法效果更优。例如,在英语与阿拉伯语之间的词性标注任务中,重对齐能使distilmBERT的准确率提升15.8个百分点,甚至超过XLM-R Large模型1.7个百分点。因此,我们主张将小型多语言模型的重对齐方法研究作为模型规模扩展之外的替代方案。