We often verbally express emotions in a multifaceted manner, they may vary in their intensities and may be expressed not just as a single but as a mixture of emotions. This wide spectrum of emotions is well-studied in the structural model of emotions, which represents variety of emotions as derivative products of primary emotions with varying degrees of intensity. In this paper, we propose an emotional text-to-speech design to simulate a wider spectrum of emotions grounded on the structural model. Our proposed design, Daisy-TTS, incorporates a prosody encoder to learn emotionally-separable prosody embedding as a proxy for emotion. This emotion representation allows the model to simulate: (1) Primary emotions, as learned from the training samples, (2) Secondary emotions, as a mixture of primary emotions, (3) Intensity-level, by scaling the emotion embedding, and (4) Emotions polarity, by negating the emotion embedding. Through a series of perceptual evaluations, Daisy-TTS demonstrated overall higher emotional speech naturalness and emotion perceiveability compared to the baseline.
翻译:我们通常以多维度方式口头表达情感,其强度可能不同,且可能以单一情感或混合情感的形式呈现。这种宽泛的情感频谱在情感结构模型中得到了充分研究,该模型将多种情感描述为具有不同强度的基础情感的衍生产物。本文提出一种基于情感结构模型的情感文本转语音设计方案,用于模拟更宽泛的情感频谱。我们提出的设计Daisy-TTS集成了韵律编码器,通过学习可分离情感的韵律嵌入作为情感的代理表示。这种情感表征使模型能够模拟:(1)基础情感(从训练样本中学习)、(2)衍生情感(作为基础情感的混合)、(3)强度层级(通过缩放情感嵌入实现)、(4)情感极性(通过取反情感嵌入实现)。通过一系列感知评估,Daisy-TTS在情感语音自然度和情感可感知度方面均显著优于基线方法。