Natural language processing of Low-Resource Languages (LRL) is often challenged by the lack of data. Therefore, achieving accurate machine translation (MT) in a low-resource environment is a real problem that requires practical solutions. Research in multilingual models have shown that some LRLs can be handled with such models. However, their large size and computational needs make their use in constrained environments (e.g., mobile/IoT devices or limited/old servers) impractical. In this paper, we address this problem by leveraging the power of large multilingual MT models using knowledge distillation. Knowledge distillation can transfer knowledge from a large and complex teacher model to a simpler and smaller student model without losing much in performance. We also make use of high-resource languages that are related or share the same linguistic root as the target LRL. For our evaluation, we consider Luxembourgish as the LRL that shares some roots and properties with German. We build multiple resource-efficient models based on German, knowledge distillation from the multilingual No Language Left Behind (NLLB) model, and pseudo-translation. We find that our efficient models are more than 30\% faster and perform only 4\% lower compared to the large state-of-the-art NLLB model.
翻译:低资源语言的自然语言处理常因数据匮乏而面临挑战。因此,在低资源环境下实现准确的机器翻译是一个需要实际解决方案的现实问题。多语言模型的研究表明,某些低资源语言可通过此类模型处理。然而,其庞大的规模和计算需求使其在受限环境(如移动/物联网设备或老旧/性能有限的服务器)中的使用不切实际。本文通过利用知识蒸馏技术来应对这一问题:知识蒸馏能将知识从大型复杂教师模型迁移至更简单紧凑的学生模型,且性能损失较小。我们还利用了与目标低资源语言相关或共享同一语言根源的高资源语言。在评估中,我们将卢森堡语作为研究对象——它与德语共享部分语言根源与特性。我们基于德语、来自多语言NLLB模型的知识蒸馏以及伪翻译,构建了多种资源高效模型。实验表明,与大规模先进NLLB模型相比,我们的高效模型速度提升超过30%,性能仅下降4%。