Consistency regularization methods, such as R-Drop (Liang et al., 2021) and CrossConST (Gao et al., 2023), have achieved impressive supervised and zero-shot performance in the neural machine translation (NMT) field. Can we also boost end-to-end (E2E) speech-to-text translation (ST) by leveraging consistency regularization? In this paper, we conduct empirical studies on intra-modal and cross-modal consistency and propose two training strategies, SimRegCR and SimZeroCR, for E2E ST in regular and zero-shot scenarios. Experiments on the MuST-C benchmark show that our approaches achieve state-of-the-art (SOTA) performance in most translation directions. The analyses prove that regularization brought by the intra-modal consistency, instead of modality gap, is crucial for the regular E2E ST, and the cross-modal consistency could close the modality gap and boost the zero-shot E2E ST performance.
翻译:一致性正则化方法(如R-Drop (Liang et al., 2021) 和 CrossConST (Gao et al., 2023))已在神经机器翻译(NMT)领域取得了显著的监督学习和零样本性能。我们能否通过利用一致性正则化来提升端到端(E2E)语音到文本翻译(ST)的性能?本文针对模态内一致性和跨模态一致性开展实证研究,并提出了两种训练策略:用于常规场景的SimRegCR和用于零样本场景的SimZeroCR。在MuST-C基准测试上的实验表明,我们的方法在大多数翻译方向上达到了当前最佳性能(SOTA)。分析证明,模态内一致性带来的正则化(而非模态差距)对常规端到端ST至关重要,而跨模态一致性能够缩小模态差距并提升零样本端到端ST的性能。