There has been significant progress in emotional Text-To-Speech (TTS) synthesis technology in recent years. However, existing methods primarily focus on the synthesis of a limited number of emotion types and have achieved unsatisfactory performance in intensity control. To address these limitations, we propose EmoMix, which can generate emotional speech with specified intensity or a mixture of emotions. Specifically, EmoMix is a controllable emotional TTS model based on a diffusion probabilistic model and a pre-trained speech emotion recognition (SER) model used to extract emotion embedding. Mixed emotion synthesis is achieved by combining the noises predicted by diffusion model conditioned on different emotions during only one sampling process at the run-time. We further apply the Neutral and specific primary emotion mixed in varying degrees to control intensity. Experimental results validate the effectiveness of EmoMix for synthesizing mixed emotion and intensity control.
翻译:近年来,情感文本到语音合成技术取得了显著进展。然而,现有方法主要聚焦于有限情感类型的合成,且在强度控制方面表现不佳。为解决这些局限,我们提出EmoMix,它能够生成具有指定强度或混合情感的语音。具体而言,EmoMix是一种基于扩散概率模型的可控情感TTS模型,并采用预训练的语音情感识别模型来提取情感嵌入。通过在运行时单次采样过程中结合由不同情感条件扩散模型预测的噪声,即可实现混合情感合成。我们进一步应用不同比例的中性情感与特定主导情感混合,以控制情感强度。实验结果验证了EmoMix在合成混合情感与强度控制方面的有效性。