Although the Transformer is currently the best-performing architecture in the homogeneous configuration (self-attention only) in Neural Machine Translation, many State-of-the-Art models in Natural Language Processing are made of a combination of different Deep Learning approaches. However, these models often focus on combining a couple of techniques only and it is unclear why some methods are chosen over others. In this work, we investigate the effectiveness of integrating an increasing number of heterogeneous methods. Based on a simple combination strategy and performance-driven synergy criteria, we designed the Multi-Encoder Transformer, which consists of up to five diverse encoders. Results showcased that our approach can improve the quality of the translation across a variety of languages and dataset sizes and it is particularly effective in low-resource languages where we observed a maximum increase of 7.16 BLEU compared to the single-encoder model.
翻译:尽管Transformer在神经机器翻译的同质配置(仅自注意力)中当前表现最佳,但自然语言处理领域的许多最先进模型由不同深度学习方法的组合构成。然而,这些模型往往仅侧重于结合少数几种技术,且部分方法被选用的原因尚不明确。本研究探究了整合不断增加数量的异构方法的有效性。基于简单组合策略与性能驱动的协同准则,我们设计了多编码器Transformer,该架构包含最多五种不同的编码器。结果表明,我们的方法能够在多种语言及数据集规模下提升翻译质量,尤其在低资源语言中效果显著——与单编码器模型相比,BLEU值最高提升7.16点。