The spontaneous behavior that often occurs in conversations makes speech more human-like compared to reading-style. However, synthesizing spontaneous-style speech is challenging due to the lack of high-quality spontaneous datasets and the high cost of labeling spontaneous behavior. In this paper, we propose a semi-supervised pre-training method to increase the amount of spontaneous-style speech and spontaneous behavioral labels. In the process of semi-supervised learning, both text and speech information are considered for detecting spontaneous behaviors labels in speech. Moreover, a linguistic-aware encoder is used to model the relationship between each sentence in the conversation. Experimental results indicate that our proposed method achieves superior expressive speech synthesis performance with the ability to model spontaneous behavior in spontaneous-style speech and predict reasonable spontaneous behavior from text.
翻译:对话中常见的自发性行为使得语音相较于朗读风格更具人类自然感。然而,由于缺乏高质量的自发性数据集以及标注自发性行为的高昂成本,合成自发风格语音具有挑战性。本文提出了一种半监督预训练方法,以增加自发性风格语音样本及自发性行为标签的数量。在半监督学习过程中,同时利用文本和语音信息检测语音中的自发性行为标签。此外,采用语言感知编码器对对话中各语句间的关系进行建模。实验结果表明,所提方法能够对自发风格语音中的自发性行为进行建模,并从文本中预测合理的自发性行为,从而实现优越的表现力语音合成性能。