This paper investigates the performance of massively multilingual neural machine translation (NMT) systems in translating Yor\`ub\'a greetings ($\varepsilon$ k\'u [MASK]), which are a big part of Yor\`ub\'a language and culture, into English. To evaluate these models, we present IkiniYor\`ub\'a, a Yor\`ub\'a-English translation dataset containing some Yor\`ub\'a greetings, and sample use cases. We analysed the performance of different multilingual NMT systems including Google and NLLB and show that these models struggle to accurately translate Yor\`ub\'a greetings into English. In addition, we trained a Yor\`ub\'a-English model by finetuning an existing NMT model on the training split of IkiniYor\`ub\'a and this achieved better performance when compared to the pre-trained multilingual NMT models, although they were trained on a large volume of data.
翻译:摘要:本文研究了大规模多语言神经机器翻译(NMT)系统在将约鲁巴问候语($\varepsilon$ k\'u [MASK],在约鲁巴语言与文化中占据重要地位)翻译成英语时的表现。为评估这些模型,我们构建了IkiniYorùbá数据集——一个包含若干约鲁巴问候语及示例用法的约鲁巴-英语翻译数据集。我们分析了包括Google和NLLB在内的多种多语言NMT系统的性能,结果表明这些模型难以准确地将约鲁巴问候语翻译成英语。此外,我们通过在IkiniYorùbá训练集上微调现有NMT模型,训练了一个约鲁巴-英语翻译模型,该模型相较于预训练的多语言NMT模型(尽管后者基于大量数据训练)取得了更优的性能。