While many parallel corpora are not publicly accessible for data copyright, data privacy and competitive differentiation reasons, trained translation models are increasingly available on open platforms. In this work, we propose a method called continual knowledge distillation to take advantage of existing translation models to improve one model of interest. The basic idea is to sequentially transfer knowledge from each trained model to the distilled model. Extensive experiments on Chinese-English and German-English datasets show that our method achieves significant and consistent improvements over strong baselines under both homogeneous and heterogeneous trained model settings and is robust to malicious models.
翻译:由于数据版权、数据隐私和竞争差异化等原因,许多平行语料库无法公开获取,而经过训练的翻译模型在开放平台上日益普及。本文提出一种名为“持续知识蒸馏”的方法,旨在利用现有翻译模型来改进特定目标模型。其基本思想是依次将每个已训练模型的知识迁移至蒸馏模型。在中英和德英数据集上的大量实验表明,该方法在同构和异构已训练模型设置下均能显著且持续地超越强基线,并对恶意模型具有鲁棒性。