Massively multilingual pre-trained language models (MMPLMs) are developed in recent years demonstrating superpowers and the pre-knowledge they acquire for downstream tasks. This work investigates whether MMPLMs can be applied to clinical domain machine translation (MT) towards entirely unseen languages via transfer learning. We carry out an experimental investigation using Meta-AI's MMPLMs ``wmt21-dense-24-wide-en-X and X-en (WMT21fb)'' which were pre-trained on 7 language pairs and 14 translation directions including English to Czech, German, Hausa, Icelandic, Japanese, Russian, and Chinese, and the opposite direction. We fine-tune these MMPLMs towards English-\textit{Spanish} language pair which \textit{did not exist at all} in their original pre-trained corpora both implicitly and explicitly. We prepare carefully aligned \textit{clinical} domain data for this fine-tuning, which is different from their original mixed domain knowledge. Our experimental result shows that the fine-tuning is very successful using just 250k well-aligned in-domain EN-ES segments for three sub-task translation testings: clinical cases, clinical terms, and ontology concepts. It achieves very close evaluation scores to another MMPLM NLLB from Meta-AI, which included Spanish as a high-resource setting in the pre-training. To the best of our knowledge, this is the first work on using MMPLMs towards \textit{clinical domain transfer-learning NMT} successfully for totally unseen languages during pre-training.
翻译:近年来,海量多语言预训练语言模型(MMPLMs)被开发出来,展现出强大的能力及其为下游任务准备的先验知识。本研究探究MMPLMs能否通过迁移学习应用于临床领域机器翻译,并翻译至预训练中完全未见的语言。我们使用Meta-AI的MMPLMs“wmt21-dense-24-wide-en-X和X-en(WMT21fb)”开展实验研究,该模型在7个语言对和14个翻译方向(包括英语与捷克语、德语、豪萨语、冰岛语、日语、俄语和汉语之间的双向翻译)上进行了预训练。我们针对英语-西班牙语语言对(该语言对在其原始预训练语料中既未隐式也未显式存在)对这些MMPLMs进行微调。为此微调过程,我们精心准备了与原始混合领域知识不同的、严格对齐的临床领域数据。实验结果表明,仅使用25万个良好对齐的领域内英-西片段,即可在三个子任务翻译测试(临床病例、临床术语和本体概念)中取得非常成功的微调效果。其评估得分与Meta-AI的另一个MMPLM模型NLLB(在预训练中已将西班牙语作为高资源设置纳入)非常接近。据我们所知,这是首次利用MMPLMs成功实现预训练中完全未见语言的临床领域迁移学习神经机器翻译的研究。