Humans often speak in a continuous manner which leads to coherent and consistent prosody properties across neighboring utterances. However, most state-of-the-art speech synthesis systems only consider the information within each sentence and ignore the contextual semantic and acoustic features. This makes it inadequate to generate high-quality paragraph-level speech which requires high expressiveness and naturalness. To synthesize natural and expressive speech for a paragraph, a context-aware speech synthesis system named MaskedSpeech is proposed in this paper, which considers both contextual semantic and acoustic features. Inspired by the masking strategy in the speech editing research, the acoustic features of the current sentence are masked out and concatenated with those of contextual speech, and further used as additional model input. The phoneme encoder takes the concatenated phoneme sequence from neighboring sentences as input and learns fine-grained semantic information from contextual text. Furthermore, cross-utterance coarse-grained semantic features are employed to improve the prosody generation. The model is trained to reconstruct the masked acoustic features with the augmentation of both the contextual semantic and acoustic features. Experimental results demonstrate that the proposed MaskedSpeech outperformed the baseline system significantly in terms of naturalness and expressiveness.
翻译:人类通常以连续方式说话,这使相邻话语间的韵律特性保持连贯一致。然而,现有最先进的语音合成系统大多仅考虑单句内部信息,而忽略了上下文语义及声学特征。这使得系统难以生成需要高表现力与自然度的段落级优质语音。为合成自然且富有表现力的段落语音,本文提出一种名为MaskedSpeech的上下文感知语音合成系统,该系统同时整合了上下文语义与声学特征。受语音编辑研究中掩码策略的启发,当前句子的声学特征被掩码后与上下文语音特征拼接,并作为额外模型输入。音素编码器将相邻句子的拼接音素序列作为输入,从上下文文本中学习细粒度语义信息。此外,跨语句的粗粒度语义特征被用于改善韵律生成。模型通过结合上下文语义与声学特征进行增强,以重构被掩码的声学特征为训练目标。实验结果表明,所提出的MaskedSpeech系统在自然度与表现力方面显著优于基线系统。