Previous pitch-controllable text-to-speech (TTS) models rely on directly modeling fundamental frequency, leading to low variance in synthesized speech. To address this issue, we propose PITS, an end-to-end pitch-controllable TTS model that utilizes variational inference to model pitch. Based on VITS, PITS incorporates the Yingram encoder, the Yingram decoder, and adversarial training of pitch-shifted synthesis to achieve pitch-controllability. Experiments demonstrate that PITS generates high-quality speech that is indistinguishable from ground truth speech and has high pitch-controllability without quality degradation. Code and audio samples will be available at https://github.com/anonymous-pits/pits.
翻译:先前的音高可控文本转语音(TTS)模型依赖于直接建模基频,导致合成语音的方差较低。为解决这一问题,我们提出PITS——一种利用变分推断对音高进行建模的端到端音高可控TTS模型。基于VITS,PITS引入Yingram编码器、Yingram解码器以及音高偏移合成的对抗训练,以实现音高可控性。实验表明,PITS能生成与真实语音难以区分的高质量语音,且在不降低质量的前提下具备高度音高可控性。代码与音频样本将在https://github.com/anonymous-pits/pits 上发布。