Existing fine-grained intensity regulation methods rely on explicit control through predicted emotion probabilities. However, these high-level semantic probabilities are often inaccurate and unsmooth at the phoneme level, leading to bias in learning. Especially when we attempt to mix multiple emotion intensities for specific phonemes, resulting in markedly reduced controllability and naturalness of the synthesis. To address this issue, we propose the CAScaded Explicit and Implicit coNtrol framework (CASEIN), which leverages accurate disentanglement of emotion manifolds from the reference speech to learn the implicit representation at a lower semantic level. This representation bridges the semantical gap between explicit probabilities and the synthesis model, reducing bias in learning. In experiments, our CASEIN surpasses existing methods in both controllability and naturalness. Notably, we are the first to achieve fine-grained control over the mixed intensity of multiple emotions.
翻译:现有细粒度强度调节方法依赖通过预测情感概率进行的显式控制。然而,这些高层语义概率在音素层面通常不准确且不光滑,导致学习偏差。特别是当尝试为特定音素混合多种情感强度时,合成结果的可控性和自然度显著下降。为解决此问题,我们提出级联显式与隐式联合控制框架(CASEIN),该框架利用从参考语音中精确解耦的情感流形,在较低语义层面学习隐式表征。该表征弥合了显式概率与合成模型之间的语义鸿沟,减少了学习偏差。实验表明,我们的CASEIN在可控性和自然度上均超越现有方法。值得注意的是,我们是首个实现对多种情感混合强度进行细粒度控制的工作。