Multimodal machine translation (MMT) aims to improve translation quality by incorporating information from other modalities, such as vision. Previous MMT systems mainly focus on better access and use of visual information and tend to validate their methods on image-related datasets. These studies face two challenges. First, they can only utilize triple data (bilingual texts with images), which is scarce; second, current benchmarks are relatively restricted and do not correspond to realistic scenarios. Therefore, this paper correspondingly establishes new methods and new datasets for MMT. First, we propose a framework 2/3-Triplet with two new approaches to enhance MMT by utilizing large-scale non-triple data: monolingual image-text data and parallel text-only data. Second, we construct an English-Chinese {e}-commercial {m}ulti{m}odal {t}ranslation dataset (including training and testing), named EMMT, where its test set is carefully selected as some words are ambiguous and shall be translated mistakenly without the help of images. Experiments show that our method is more suitable for real-world scenarios and can significantly improve translation performance by using more non-triple data. In addition, our model also rivals various SOTA models in conventional multimodal translation benchmarks.
翻译:多模态机器翻译旨在通过整合其他模态(如视觉)的信息来提升翻译质量。以往的多模态机器翻译系统主要侧重于更好地获取和利用视觉信息,并倾向于在图像相关数据集上验证其方法。这些研究面临两大挑战:第一,它们只能利用三元组数据(带图像的双语文本),而此类数据稀缺;第二,当前的基准测试相对受限,且不符合实际场景。因此,本文相应地提出了新的多模态机器翻译方法和数据集。首先,我们提出了一个名为“2/3-三元组”的框架,通过两种新方法利用大规模非三元组数据(单语言图像文本数据和平行纯文本数据)来增强多模态机器翻译。其次,我们构建了一个包含训练集和测试集的英汉电商多模态翻译数据集,命名为EMMT,其测试集经过精心挑选:部分单词具有歧义性,若无图像辅助则容易误译。实验表明,我们的方法更适用于真实场景,且通过利用更多非三元组数据可显著提升翻译性能。此外,我们的模型在传统多模态翻译基准测试中亦能与多种先进模型相匹敌。