Multilingual sequence-to-sequence models perform poorly with increased language coverage and fail to consistently generate text in the correct target language in few-shot settings. To address these challenges, we propose mmT5, a modular multilingual sequence-to-sequence model. mmT5 utilizes language-specific modules during pre-training, which disentangle language-specific information from language-agnostic information. We identify representation drift during fine-tuning as a key limitation of modular generative models and develop strategies that enable effective zero-shot transfer. Our model outperforms mT5 at the same parameter sizes by a large margin on representative natural language understanding and generation tasks in 40+ languages. Compared to mT5, mmT5 raises the rate of generating text in the correct language under zero-shot settings from 7% to 99%, thereby greatly alleviating the source language hallucination problem.
翻译:多语言序列到序列模型在语言覆盖范围扩大时表现不佳,并且在少样本设置下无法一致地生成正确目标语言的文本。为了解决这些挑战,我们提出了mmT5,一种模块化的多语言序列到序列模型。mmT5在预训练期间使用语言专用模块,将语言特定信息与语言无关信息分离。我们识别出微调过程中的表示漂移是模块化生成模型的一个关键限制,并开发了能够实现有效零样本迁移的策略。我们的模型在40多种语言的代表性自然语言理解和生成任务上,以相同的参数规模大幅超越mT5。与mT5相比,mmT5在零样本设置下生成正确语言文本的比率从7%提升至99%,从而极大缓解了源语言幻觉问题。