Multilingual language models have shown impressive cross-lingual transfer ability across a diverse set of languages and tasks. To improve the cross-lingual ability of these models, some strategies include transliteration and finer-grained segmentation into characters as opposed to subwords. In this work, we investigate lexical sharing in multilingual machine translation (MT) from Hindi, Gujarati, Nepali into English. We explore the trade-offs that exist in translation performance between data sampling and vocabulary size, and we explore whether transliteration is useful in encouraging cross-script generalisation. We also verify how the different settings generalise to unseen languages (Marathi and Bengali). We find that transliteration does not give pronounced improvements and our analysis suggests that our multilingual MT models trained on original scripts seem to already be robust to cross-script differences even for relatively low-resource languages
翻译:多语言语言模型在多种语言和任务中展现了显著的跨语言迁移能力。为提升这些模型的跨语言能力,一些策略包括音译以及将子词分割细化为字符级分词。在本工作中,我们研究了从印地语、古吉拉特语、尼泊尔语到英语的多语言机器翻译中的词汇共享现象。我们探讨了数据采样与词汇量之间对翻译性能的权衡,并验证了音译是否有助于促进跨书写系统的泛化。此外,我们还检验了不同设置对未见语言(马拉地语和孟加拉语)的泛化效果。研究发现,音译并未带来显著改进,我们的分析表明,基于原始书写系统训练的多语言机器翻译模型似乎已能有效应对跨书写系统的差异,即使对于资源相对匮乏的语言也是如此。