Recent developments in neural speech synthesis and vocoding have sparked a renewed interest in voice conversion (VC). Beyond timbre transfer, achieving controllability on para-linguistic parameters such as pitch and Speed is critical in deploying VC systems in many application scenarios. Existing studies, however, either only provide utterance-level global control or lack interpretability on the controls. In this paper, we propose ControlVC, the first neural voice conversion system that achieves time-varying controls on pitch and speed. ControlVC uses pre-trained encoders to compute pitch and linguistic embeddings from the source utterance and speaker embeddings from the target utterance. These embeddings are then concatenated and converted to speech using a vocoder. It achieves speed control through TD-PSOLA pre-processing on the source utterance, and achieves pitch control by manipulating the pitch contour before feeding it to the pitch encoder. Systematic subjective and objective evaluations are conducted to assess the speech quality and controllability. Results show that, on non-parallel and zero-shot conversion tasks, ControlVC significantly outperforms two other self-constructed baselines on speech quality, and it can successfully achieve time-varying pitch and speed control.
翻译:近年来,神经语音合成与声码器技术的进步重新激发了人们对语音转换(VC)的研究兴趣。除音色迁移外,对音高、速度等副语言参数实现可控性,是语音转换系统在众多应用场景中部署的关键。然而,现有研究要么仅提供语句级别的全局控制,要么缺乏对控制机制的可解释性。本文提出ControlVC,这是首个实现音高与速度时变控制的神经语音转换系统。ControlVC采用预训练编码器从源语音中提取音高与语言嵌入向量,并从目标语音中提取说话人嵌入向量,随后将这些嵌入向量拼接后通过声码器转换为语音。该系统通过TD-PSOLA预处理对源语音实现速度控制,并通过在输入音高编码器前对音高轮廓进行操控来实现音高控制。我们通过系统的客观与主观评估对语音质量与可控性进行评测。结果表明,在非平行与零样本转换任务中,ControlVC在语音质量上显著优于另外两个自建基线系统,并能成功实现音高与速度的时变控制。