Nowadays, as more and more systems achieve good performance in traditional voice conversion (VC) tasks, people's attention gradually turns to VC tasks under extreme conditions. In this paper, we propose a novel method for zero-shot voice conversion. We aim to obtain intermediate representations for speaker-content disentanglement of speech to better remove speaker information and get pure content information. Accordingly, our proposed framework contains a module that removes the speaker information from the acoustic feature of the source speaker. Moreover, speaker information control is added to our system to maintain the voice cloning performance. The proposed system is evaluated by subjective and objective metrics. Results show that our proposed system significantly reduces the trade-off problem in zero-shot voice conversion, while it also manages to have high spoofing power to the speaker verification system.
翻译:摘要:当前,随着越来越多系统在传统语音转换任务中取得优异性能,研究关注点逐渐转向极端条件下的语音转换任务。本文提出了一种新颖的零样本语音转换方法,旨在获取语音中说话人与内容解耦的中间表征,以更有效地去除说话人信息并提取纯净内容信息。据此,所提框架包含一个模块,用于从源说话人的声学特征中移除说话人信息。此外,系统增加了说话人信息控制机制以维持语音克隆性能。通过主客观指标评估,结果表明所提系统显著缓解了零样本语音转换中的权衡问题,同时对说话人验证系统展现出高欺骗能力。