This paper presents Translatotron 3, a novel approach to train a direct speech-to-speech translation model from monolingual speech-text datasets only in a fully unsupervised manner. Translatotron 3 combines masked autoencoder, unsupervised embedding mapping, and back-translation to achieve this goal. Experimental results in speech-to-speech translation tasks between Spanish and English show that Translatotron 3 outperforms a baseline cascade system, reporting 18.14 BLEU points improvement on the synthesized Unpaired-Conversational dataset. In contrast to supervised approaches that necessitate real paired data, which is unavailable, or specialized modeling to replicate para-/non-linguistic information, Translatotron 3 showcases its capability to retain para-/non-linguistic such as pauses, speaking rates, and speaker identity. Audio samples can be found in our website http://google-research.github.io/lingvo-lab/translatotron3
翻译:本文提出Translatotron 3,一种完全无监督地从单语语音-文本数据集训练直接语音到语音翻译模型的新方法。Translatotron 3结合了掩码自编码器、无监督嵌入映射和反向翻译以实现这一目标。在西班牙语与英语之间的语音到语音翻译任务中的实验结果表明,Translatotron 3优于基线级联系统,在合成的Unpaired-Conversational数据集上报告了18.14个BLEU分数的提升。与需要真实配对数据(现实中不可获得)或专门建模以复现副语言/非语言信息的监督方法不同,Translatotron 3展示了其保留诸如停顿、语速和说话人身份等副语言/非语言信息的能力。音频样本可在我们的网站http://google-research.github.io/lingvo-lab/translatotron3上获取。