In this paper, we present MuLanTTS, the Microsoft end-to-end neural text-to-speech (TTS) system designed for the Blizzard Challenge 2023. About 50 hours of audiobook corpus for French TTS as hub task and another 2 hours of speaker adaptation as spoke task are released to build synthesized voices for different test purposes including sentences, paragraphs, homographs, lists, etc. Building upon DelightfulTTS, we adopt contextual and emotion encoders to adapt the audiobook data to enrich beyond sentences for long-form prosody and dialogue expressiveness. Regarding the recording quality, we also apply denoise algorithms and long audio processing for both corpora. For the hub task, only the 50-hour single speaker data is used for building the TTS system, while for the spoke task, a multi-speaker source model is used for target speaker fine tuning. MuLanTTS achieves mean scores of quality assessment 4.3 and 4.5 in the respective tasks, statistically comparable with natural speech while keeping good similarity according to similarity assessment. The excellent quality and similarity in this year's new and dense statistical evaluation.
翻译:本文介绍了MuLanTTS,即微软为2023年Blizzard挑战赛设计的端到端神经文本转语音(TTS)系统。为构建用于不同测试目的(包括句子、段落、同形异义词、列表等)的合成语音,系统发布了约50小时的法语TTS有声读物语料库作为中心任务,以及另外2小时的说话人自适应语料库作为分支任务。在DelightfulTTS的基础上,我们采用上下文和情感编码器对有声读物数据进行适配,以增强超越句子级别的长时韵律和对话表现力。针对录音质量,我们还对两个语料库应用了去噪算法和长音频处理。中心任务仅使用50小时单说话人数据构建TTS系统,而分支任务则采用多说话人源模型进行目标说话人微调。在各自任务中,MuLanTTS的质量评估平均得分分别为4.3和4.5,在统计上与自然语音相当,同时根据相似度评估保持了良好的相似性。这一优异的质量与相似性在今年全新且密集的统计评估中得以体现。