This paper presents Translatotron 3, a novel approach to train a direct speech-to-speech translation model from monolingual speech-text datasets only in a fully unsupervised manner. Translatotron 3 combines masked autoencoder, unsupervised embedding mapping, and back-translation to achieve this goal. Experimental results in speech-to-speech translation tasks between Spanish and English show that Translatotron 3 outperforms a baseline cascade system, reporting 18.14 BLEU points improvement on the synthesized Unpaired-Conversational dataset. In contrast to supervised approaches that necessitate real paired data, which is unavailable, or specialized modeling to replicate para-/non-linguistic information, Translatotron 3 showcases its capability to retain para-/non-linguistic such as pauses, speaking rates, and speaker identity. Audio samples can be found in our website http://google-research.github.io/lingvo-lab/translatotron3
翻译:本文提出Translatotron 3,这是一种全新的方法,它仅利用单语语音-文本数据集,以完全无监督的方式训练直接的语音到语音翻译模型。Translatotron 3结合了掩码自编码器、无监督嵌入映射和反向翻译以实现这一目标。在西班牙语和英语之间的语音到语音翻译任务中,实验结果表明,Translatotron 3在合成的Unpaired-Conversational数据集上超越了基线级联系统,BLEU值提升了18.14分。与需要真实配对数据(通常不可获得)或专门建模以复制副语言/非语言信息的有监督方法相比,Translatotron 3展示了其保留副语言/非语言信息(如停顿、语速和说话人身份)的能力。音频样本可在我们的网站http://google-research.github.io/lingvo-lab/translatotron3上获取。