In this paper, we propose a differentiable WORLD synthesizer and demonstrate its use in end-to-end audio style transfer tasks such as (singing) voice conversion and the DDSP timbre transfer task. Accordingly, our baseline differentiable synthesizer has no model parameters, yet it yields adequate synthesis quality. We can extend the baseline synthesizer by appending lightweight black-box postnets which apply further processing to the baseline output in order to improve fidelity. An alternative differentiable approach considers extraction of the source excitation spectrum directly, which can improve naturalness albeit for a narrower class of style transfer applications. The acoustic feature parameterization used by our approaches has the added benefit that it naturally disentangles pitch and timbral information so that they can be modeled separately. Moreover, as there exists a robust means of estimating these acoustic features from monophonic audio sources, it allows for parameter loss terms to be added to an end-to-end objective function, which can help convergence and/or further stabilize (adversarial) training.
翻译:本文提出一种可微WORLD合成器,并展示其在端到端音频风格迁移任务中的应用,包括(歌唱)语音转换和DDSP音色迁移任务。我们的基线可微合成器不含模型参数,却能产生足够好的合成质量。我们可通过在基线合成器后附加轻量级黑盒后置网络来进一步处理基线输出,从而提升保真度。另一种可微方法则考虑直接提取声源激励频谱,虽适用于较窄范围的风格迁移应用,但可改善自然度。本文方法采用的声学特征参数化方案具有额外优势:它天然地将音高与音色信息解耦,便于分别建模。此外,由于存在从单声道音频源中稳健估计这些声学特征的方法,我们可在端到端目标函数中添加参数损失项,这有助于加速收敛和/或进一步稳定(对抗)训练。