Multilingual models have been widely used for cross-lingual transfer to low-resource languages. However, the performance on these languages is hindered by their underrepresentation in the pretraining data. To alleviate this problem, we propose a novel multilingual training technique based on teacher-student knowledge distillation. In this setting, we utilize monolingual teacher models optimized for their language. We use those teachers along with balanced (sub-sampled) data to distill the teachers' knowledge into a single multilingual student. Our method outperforms standard training methods in low-resource languages and retrains performance on high-resource languages while using the same amount of data. If applied widely, our approach can increase the representation of low-resource languages in NLP systems.
翻译:多语言模型已被广泛用于向低资源语言的跨语言迁移。然而,预训练数据中低资源语言代表性不足,阻碍了这些语言上的表现。为缓解该问题,我们提出一种基于师生知识蒸馏的新型多语言训练技术。在该设定下,我们利用针对各自语言优化的单语教师模型,并结合平衡(子采样)数据,将教师知识蒸馏至单个多语言学生模型。我们的方法在低资源语言上优于标准训练方法,同时在使用等量数据的情况下,保持了高资源语言上的性能。若广泛采用,该方法可提升低资源语言在NLP系统中的代表性。