We propose SelfVC, a training strategy to iteratively improve a voice conversion model with self-synthesized examples. Previous efforts on voice conversion focus on explicitly disentangling speech representations to separately encode speaker characteristics and linguistic content. However, disentangling speech representations to capture such attributes using task-specific loss terms can lead to information loss by discarding finer nuances of the original signal. In this work, instead of explicitly disentangling attributes with loss terms, we present a framework to train a controllable voice conversion model on entangled speech representations derived from self-supervised learning and speaker verification models. First, we develop techniques to derive prosodic information from the audio signal and SSL representations to train predictive submodules in the synthesis model. Next, we propose a training strategy to iteratively improve the synthesis model for voice conversion, by creating a challenging training objective using self-synthesized examples. In this training approach, the current state of the synthesis model is used to generate voice-converted variations of an utterance, which serve as inputs for the reconstruction task, ensuring a continuous and purposeful refinement of the model. We demonstrate that incorporating such self-synthesized examples during training improves the speaker similarity of generated speech as compared to a baseline voice conversion model trained solely on heuristically perturbed inputs. SelfVC is trained without any text and is applicable to a range of tasks such as zero-shot voice conversion, cross-lingual voice conversion, and controllable speech synthesis with pitch and pace modifications. SelfVC achieves state-of-the-art results in zero-shot voice conversion on metrics evaluating naturalness, speaker similarity, and intelligibility of synthesized audio.
翻译:我们提出SelfVC,一种通过自我合成示例迭代改进语音转换模型的训练策略。以往语音转换研究主要聚焦于明确解耦语音表征,以分别编码说话人特征与语言内容。然而,采用任务特定损失项捕捉此类属性的解耦方法可能因舍弃原始信号的细微差异而导致信息损失。本研究提出一个框架,无需显式解耦属性的损失项,即可基于自监督学习与说话人验证模型推导的耦合语音表征训练可控语音转换模型。首先,我们开发从音频信号和SSL表征中提取韵律信息的技术,以训练合成模型中的预测子模块。其次,提出一种通过构建具有挑战性的自我合成示例训练目标,迭代改进语音转换合成模型的策略。在该训练方法中,利用合成模型的当前状态生成语音转换变体,将其作为重构任务的输入,确保模型持续且方向明确的优化。实验表明,与仅采用启发式扰动输入训练的基线语音转换模型相比,在训练中引入此类自我合成示例可提升生成语音的说话人相似度。SelfVC无需文本即可训练,适用于零样本语音转换、跨语言语音转换及具备音高与语速控制的可控语音合成等任务。在自然度、说话人相似度与可懂度指标上,SelfVC在零样本语音转换任务中达到了当前最优水平。