We propose a cross-lingual neural codec language model, VALL-E X, for cross-lingual speech synthesis. Specifically, we extend VALL-E and train a multi-lingual conditional codec language model to predict the acoustic token sequences of the target language speech by using both the source language speech and the target language text as prompts. VALL-E X inherits strong in-context learning capabilities and can be applied for zero-shot cross-lingual text-to-speech synthesis and zero-shot speech-to-speech translation tasks. Experimental results show that it can generate high-quality speech in the target language via just one speech utterance in the source language as a prompt while preserving the unseen speaker's voice, emotion, and acoustic environment. Moreover, VALL-E X effectively alleviates the foreign accent problems, which can be controlled by a language ID. Audio samples are available at \url{https://aka.ms/vallex}.
翻译:我们提出了一种跨语言神经编解码语言模型VALL-E X,用于跨语言语音合成。具体而言,我们扩展了VALL-E,训练了一个多语言条件编解码语言模型,通过使用源语言语音和目标语言文本作为提示,预测目标语言语音的声学标记序列。VALL-E X继承了强大的上下文学习能力,可应用于零样本跨语言文本到语音合成和零样本语音到语音翻译任务。实验结果表明,仅需一段源语言语音作为提示,它就能生成高质量的目标语言语音,同时保留未见说话人的声音、情感和声学环境。此外,VALL-E X有效缓解了外语口音问题,这一问题可通过语言ID进行控制。音频样本见\url{https://aka.ms/vallex}。