Neural Machine Translation (NMT) systems built on multilingual sequence-to-sequence Language Models (msLMs) fail to deliver expected results when the amount of parallel data for a language, as well as the language's representation in the model are limited. This restricts the capabilities of domain-specific NMT systems for low-resource languages (LRLs). As a solution, parallel data from auxiliary domains can be used either to fine-tune or to further pre-train the msLM. We present an evaluation of the effectiveness of these two techniques in the context of domain-specific LRL-NMT. We also explore the impact of domain divergence on NMT model performance. We recommend several strategies for utilizing auxiliary parallel data in building domain-specific NMT models for LRLs.
翻译:基于多语言序列到序列语言模型(msLM)构建的神经机器翻译(NMT)系统,在语言的平行数据量及该语言在模型中的表征有限时,无法达到预期效果。这限制了低资源语言(LRL)领域特定NMT系统的能力。作为解决方案,来自辅助领域的平行数据可用于微调或进一步预训练msLM。我们评估了这两种技术在领域特定低资源语言神经机器翻译(LRL-NMT)中的有效性,并探讨了领域差异对NMT模型性能的影响。我们推荐了几种利用辅助平行数据为低资源语言构建领域特定NMT模型的策略。