One of the major challenges of machine translation (MT) is ambiguity, which can in some cases be resolved by accompanying context such as images. However, recent work in multimodal MT (MMT) has shown that obtaining improvements from images is challenging, limited not only by the difficulty of building effective cross-modal representations, but also by the lack of specific evaluation and training data. We present a new MMT approach based on a strong text-only MT model, which uses neural adapters, a novel guided self-attention mechanism and which is jointly trained on both visually-conditioned masking and MMT. We also introduce CoMMuTE, a Contrastive Multilingual Multimodal Translation Evaluation set of ambiguous sentences and their possible translations, accompanied by disambiguating images corresponding to each translation. Our approach obtains competitive results compared to strong text-only models on standard English-to-French, English-to-German and English-to-Czech benchmarks and outperforms baselines and state-of-the-art MMT systems by a large margin on our contrastive test set. Our code and CoMMuTE are freely available.
翻译:机器翻译(MT)的主要挑战之一是歧义,在某些情况下,可以通过图像等伴随语境来消除歧义。然而,多模态机器翻译(MMT)的最新研究表明,从图像中获取改进具有挑战性,不仅受限于构建有效跨模态表示的难度,还缺乏专门的评估和训练数据。我们提出了一种基于强文本MT模型的新型MMT方法,该方法采用神经适配器、一种新颖的引导自注意力机制,并在视觉条件掩码和MMT上联合训练。我们还引入了CoMMuTE,这是一个由歧义句子及其可能译文组成的对比性多语言多模态翻译评估集,并附有对应每个译文的消歧图像。在标准英法、英德和英捷基准测试中,我们的方法相比强文本模型取得了具有竞争力的结果,并在我们的对比测试集上以较大优势超越了基线模型和现有最先进的MMT系统。我们的代码和CoMMuTE已免费公开。