We present our work on developing a multilingual, efficient text-to-text transformer that is suitable for handling long inputs. This model, called mLongT5, builds upon the architecture of LongT5, while leveraging the multilingual datasets used for pretraining mT5 and the pretraining tasks of UL2. We evaluate this model on a variety of multilingual summarization and question-answering tasks, and the results show stronger performance for mLongT5 when compared to existing multilingual models such as mBART or M-BERT.
翻译:我们提出了一个适用于处理长输入的多语言高效文本到文本Transformer模型的研究工作。该模型名为mLongT5,基于LongT5的架构构建,同时利用了mT5预训练所使用的多语言数据集以及UL2的预训练任务。我们在多种多语言摘要和问答任务上评估了该模型,结果表明,与mBART或M-BERT等现有多语言模型相比,mLongT5展现出更强的性能。