Due to Multilingual Neural Machine Translation's (MNMT) capability of zero-shot translation, many works have been carried out to fully exploit the potential of MNMT in zero-shot translation. It is often hypothesized that positional information may hinder the MNMT from outputting a robust encoded representation for decoding. However, previous approaches treat all the positional information equally and thus are unable to selectively remove certain positional information. In sharp contrast, this paper investigates how to learn to selectively preserve useful positional information. We describe the specific mechanism of positional information influencing MNMT from the perspective of linguistics at the token level. We design a token-level position disentangle module (TPDM) framework to disentangle positional information at the token level based on the explanation. Our experiments demonstrate that our framework improves zero-shot translation by a large margin while reducing the performance loss in the supervised direction compared to previous works.
翻译:基于多语言神经机器翻译(MNMT)在零样本翻译中的能力,已有大量研究致力于充分挖掘MNMT在该任务中的潜力。通常认为位置信息可能阻碍MNMT输出鲁棒的编码表示用于解码。然而,现有方法对所有位置信息同等对待,因而无法选择性移除某些位置信息。与此形成鲜明对比,本文研究了如何学习选择性保留有用的位置信息。我们从词元层面语言学角度描述了位置信息影响MNMT的具体机制,并基于该解释设计了词元级位置解耦模块(TPDM)框架,以在词元层面解耦位置信息。实验表明,与先前工作相比,本框架在显著提升零样本翻译性能的同时,减少了监督方向上的性能损失。