Multi-media communications facilitate global interaction among people. However, despite researchers exploring cross-lingual translation techniques such as machine translation and audio speech translation to overcome language barriers, there is still a shortage of cross-lingual studies on visual speech. This lack of research is mainly due to the absence of datasets containing visual speech and translated text pairs. In this paper, we present \textbf{AVMuST-TED}, the first dataset for \textbf{A}udio-\textbf{V}isual \textbf{Mu}ltilingual \textbf{S}peech \textbf{T}ranslation, derived from \textbf{TED} talks. Nonetheless, visual speech is not as distinguishable as audio speech, making it difficult to develop a mapping from source speech phonemes to the target language text. To address this issue, we propose MixSpeech, a cross-modality self-learning framework that utilizes audio speech to regularize the training of visual speech tasks. To further minimize the cross-modality gap and its impact on knowledge transfer, we suggest adopting mixed speech, which is created by interpolating audio and visual streams, along with a curriculum learning strategy to adjust the mixing ratio as needed. MixSpeech enhances speech translation in noisy environments, improving BLEU scores for four languages on AVMuST-TED by +1.4 to +4.2. Moreover, it achieves state-of-the-art performance in lip reading on CMLR (11.1\%), LRS2 (25.5\%), and LRS3 (28.0\%).
翻译:多媒体通信促进了全球范围内的人际互动。然而,尽管研究者们探索了机器翻译、音频语音翻译等跨语言翻译技术以克服语言障碍,但针对视觉语音的跨语言研究仍存在不足。这一研究空白主要源于缺乏包含视觉语音与翻译文本对的数据集。本文提出了 **AVMuST-TED**,这是首个面向**音**-**视**频**多**语言**语**音**翻**译的数据集,源自 **TED** 演讲。然而,视觉语音的辨识度低于音频语音,这使得建立源语音音素到目标语言文本的映射变得困难。为解决该问题,我们提出 MixSpeech,一种利用音频语音约束视觉语音任务训练的跨模态自学习框架。为进一步缩小跨模态差异及其对知识迁移的影响,我们建议采用通过音频流与视频流插值生成的混合语音,并结合课程学习策略动态调整混合比例。MixSpeech 在噪声环境下能有效提升语音翻译性能,在 AVMuST-TED 上四种语言的 BLEU 分数提升了 +1.4 至 +4.2。此外,它在 CMLR(11.1%)、LRS2(25.5%)和 LRS3(28.0%)数据集上的唇读任务中达到了最先进水平。