Direct speech-to-speech translation (S2ST) with discrete self-supervised representations has achieved remarkable accuracy, but is unable to preserve the speaker timbre of the source speech during translation. Meanwhile, the scarcity of high-quality speaker-parallel data poses a challenge for learning style transfer between source and target speech. We propose an S2ST framework with an acoustic language model based on discrete units from a self-supervised model and a neural codec for style transfer. The acoustic language model leverages self-supervised in-context learning, acquiring the ability for style transfer without relying on any speaker-parallel data, thereby overcoming the issue of data scarcity. By using extensive training data, our model achieves zero-shot cross-lingual style transfer on previously unseen source languages. Experiments show that our model generates translated speeches with high fidelity and style similarity. Audio samples are available at http://stylelm.github.io/ .
翻译:直接语音到语音翻译 (S2ST) 利用离散自监督表示已取得显著准确度,但无法在翻译过程中保留源语音的说话人音色。同时,高质量说话人平行数据的稀缺性给源语音与目标语音之间的风格迁移学习带来了挑战。我们提出了一种基于自监督模型离散单元的声学语言模型与神经编解码器的S2ST框架。该声学语言模型利用自监督上下文学习,无需依赖任何说话人平行数据即可获得风格迁移能力,从而克服了数据稀缺问题。通过使用大规模训练数据,我们的模型在未见过的源语言上实现了零样本跨语言风格迁移。实验表明,我们的模型生成高保真度且风格相似的目标语音。音频样本见 http://stylelm.github.io/ 。