This paper aims to synthesize the target speaker's speech with desired speaking style and emotion by transferring the style and emotion from reference speech recorded by other speakers. We address this challenging problem with a two-stage framework composed of a text-to-style-and-emotion (Text2SE) module and a style-and-emotion-to-wave (SE2Wave) module, bridging by neural bottleneck (BN) features. To further solve the multi-factor (speaker timbre, speaking style and emotion) decoupling problem, we adopt the multi-label binary vector (MBV) and mutual information (MI) minimization to respectively discretize the extracted embeddings and disentangle these highly entangled factors in both Text2SE and SE2Wave modules. Moreover, we introduce a semi-supervised training strategy to leverage data from multiple speakers, including emotion-labeled data, style-labeled data, and unlabeled data. To better transfer the fine-grained expression from references to the target speaker in non-parallel transfer, we introduce a reference-candidate pool and propose an attention-based reference selection approach. Extensive experiments demonstrate the good design of our model.
翻译:本文旨在通过迁移其他说话人参考语音中的风格和情感,合成具有目标说话人语音特性的、带有期望说话风格和情感的语音。我们提出一个两阶段框架来解决这一具有挑战性的问题,该框架由文本到风格与情感(Text2SE)模块和风格与情感到波形(SE2Wave)模块组成,两者通过神经瓶颈(BN)特征进行连接。为进一步解决多因素(说话人音色、说话风格和情感)解耦问题,我们采用多标签二元向量(MBV)和互信息(MI)最小化方法,分别在Text2SE和SE2Wave模块中对提取的嵌入进行离散化处理,并解开这些高度纠缠的因素。此外,我们引入一种半监督训练策略,以利用来自多个说话人的数据,包括情感标注数据、风格标注数据和未标注数据。为了在非平行迁移中更好地将参考语音中的细粒度表达传递给目标说话人,我们引入了一个参考候选池,并提出一种基于注意力的参考选择方法。大量实验证明了我们模型设计的优越性。