Discrete speech tokens have been more and more popular in multiple speech processing fields, including automatic speech recognition (ASR), text-to-speech (TTS) and singing voice synthesis (SVS). In this paper, we describe the systems developed by the SJTU X-LANCE group for the TTS (acoustic + vocoder), SVS, and ASR tracks in the Interspeech 2024 Speech Processing Using Discrete Speech Unit Challenge. Notably, we achieved 1st rank on the leaderboard in the TTS track both with the whole training set and only 1h training data, along with the lowest bitrate among all submissions.
翻译:离散语音标记在自动语音识别(ASR)、文本转语音(TTS)及歌声合成(SVS)等多个语音处理领域中日益普及。本文描述了上海交通大学X-LANCE团队在Interspeech 2024离散语音单元挑战赛中针对TTS(声学模型+声码器)、SVS及ASR赛道开发的系统。值得注意的是,我们在TTS赛道中,无论是使用完整训练集还是仅1小时训练数据,均取得了排行榜第一名的成绩,且在所有参赛方案中实现了最低比特率。