In this paper, we present MuLanTTS, the Microsoft end-to-end neural text-to-speech (TTS) system designed for the Blizzard Challenge 2023. About 50 hours of audiobook corpus for French TTS as hub task and another 2 hours of speaker adaptation as spoke task are released to build synthesized voices for different test purposes including sentences, paragraphs, homographs, lists, etc. Building upon DelightfulTTS, we adopt contextual and emotion encoders to adapt the audiobook data to enrich beyond sentences for long-form prosody and dialogue expressiveness. Regarding the recording quality, we also apply denoise algorithms and long audio processing for both corpora. For the hub task, only the 50-hour single speaker data is used for building the TTS system, while for the spoke task, a multi-speaker source model is used for target speaker fine tuning. MuLanTTS achieves mean scores of quality assessment 4.3 and 4.5 in the respective tasks, statistically comparable with natural speech while keeping good similarity according to similarity assessment. The excellent quality and similarity in this year's new and dense statistical evaluation.
翻译:本文介绍MuLanTTS——微软为Blizzard Challenge 2023设计的端到端神经文本转语音(TTS)系统。作为核心任务,我们发布了约50小时的法语音频书语料库用于法语音频TTS;另外还发布了2小时的说话人自适应语料库作为分支任务,用于构建面向不同测试目的(包括句子、段落、同形异义词、列表等)的合成语音。基于DelightfulTTS,我们采用上下文编码器和情感编码器对音频书数据进行适配,以增强超越句子级别的长句韵律和对话表现力。针对录音质量,我们对两个语料库均应用了降噪算法和长音频处理技术。核心任务仅使用50小时单说话人数据构建TTS系统,而分支任务则采用多说话人源模型进行目标说话人微调。在质量评估中,MuLanTTS在两个任务中分别获得4.3和4.5的平均分,在统计上与自然语音相当,同时根据相似度评估保持了良好的相似性。本年度新型密集统计评估证实了其优异的语音质量与相似度。