Transformers have achieved great success in machine translation, but transformer-based NMT models often require millions of bilingual parallel corpus for training. In this paper, we propose a novel architecture named as attention link (AL) to help improve transformer models' performance, especially in low training resources. We theoretically demonstrate the superiority of our attention link architecture in low training resources. Besides, we have done a large number of experiments, including en-de, de-en, en-fr, en-it, it-en, en-ro translation tasks on the IWSLT14 dataset as well as real low resources scene on bn-gu and gu-ta translation tasks on the CVIT PIB dataset. All the experiment results show our attention link is powerful and can lead to a significant improvement. In addition, we achieve a 37.9 BLEU score, a new sota, on the IWSLT14 de-en task by combining our attention link and other advanced methods.
翻译:Transformer在机器翻译中取得了巨大成功,但基于Transformer的神经机器翻译模型通常需要数百万双语平行语料库进行训练。本文提出了一种名为注意力链接(Attention Link, AL)的新型架构,旨在提升Transformer模型在低资源训练环境下的性能。我们从理论上证明了注意力链接架构在低训练资源中的优越性。此外,我们进行了大量实验,包括在IWSLT14数据集上的英德、德英、英法、英意、意英、英罗翻译任务,以及在CVIT PIB数据集上的孟加拉语-古吉拉特语和古吉拉特语-泰米尔语翻译任务等真实低资源场景。所有实验结果表明,我们的注意力链接架构具有强大性能,并能带来显著改进。此外,通过将注意力链接与其他先进方法相结合,我们在IWSLT14德英任务上取得了37.9 BLEU分数,达到了新的最先进水平(SOTA)。