This paper investigates the performance of massively multilingual neural machine translation (NMT) systems in translating Yor\`ub\'a greetings ($\mathcal{E}$ k\'u [MASK]), which are a big part of Yor\`ub\'a language and culture, into English. To evaluate these models, we present IkiniYor\`ub\'a, a Yor\`ub\'a-English translation dataset containing some Yor\`ub\'a greetings, and sample use cases. We analysed the performance of different multilingual NMT systems including Google and NLLB and show that these models struggle to accurately translate Yor\`ub\'a greetings into English. In addition, we trained a Yor\`ub\'a-English model by finetuning an existing NMT model on the training split of IkiniYor\`ub\'a and this achieved better performance when compared to the pre-trained multilingual NMT models, although they were trained on a large volume of data.
翻译:摘要:本文研究了大规模多语言神经机器翻译(NMT)系统在将约鲁巴语问候语($\mathcal{E}$ k\'u [MASK])——这一约鲁巴语言与文化的重要组成部分——翻译为英语时的性能表现。为评估这些模型,我们构建了IkiniYorùbá数据集,该数据集包含部分约鲁巴问候语及其使用示例。通过分析包括Google和NLLB在内的不同多语言神经机器翻译系统的性能,我们发现这些模型在将约鲁巴问候语准确翻译为英语方面存在困难。此外,我们通过在IkiniYorùbá训练集上微调现有NMT模型,训练了一个约鲁巴语-英语翻译模型。实验结果表明,该模型相较于预训练的多语言NMT模型(尽管后者基于大规模数据训练)取得了更优的翻译性能。