We propose PromptTTS++, a prompt-based text-to-speech (TTS) synthesis system that allows control over speaker identity using natural language descriptions. To control speaker identity within the prompt-based TTS framework, we introduce the concept of speaker prompt, which describes voice characteristics (e.g., gender-neutral, young, old, and muffled) designed to be approximately independent of speaking style. Since there is no large-scale dataset containing speaker prompts, we first construct a dataset based on the LibriTTS-R corpus with manually annotated speaker prompts. We then employ a diffusion-based acoustic model with mixture density networks to model diverse speaker factors in the training data. Unlike previous studies that rely on style prompts describing only a limited aspect of speaker individuality, such as pitch, speaking speed, and energy, our method utilizes an additional speaker prompt to effectively learn the mapping from natural language descriptions to the acoustic features of diverse speakers. Our subjective evaluation results show that the proposed method can better control speaker characteristics than the methods without the speaker prompt. Audio samples are available at https://reppy4620.github.io/demo.promptttspp/.
翻译:我们提出PromptTTS++,一种基于提示的文本转语音(TTS)合成系统,可通过自然语言描述控制说话者身份。为在提示型TTS框架内控制说话者身份,我们引入说话者提示概念,该提示描述语音特征(如中性性别、年轻、年老、低沉),并设计为与说话风格近似无关。由于缺乏包含说话者提示的大规模数据集,我们首先基于LibriTTS-R语料库构建数据集,并手工标注说话者提示。随后,我们采用结合混合密度网络的扩散声学模型,对训练数据中的多样说话者因素进行建模。与仅依赖描述音高、语速和能量等有限说话者个性方面的风格提示的现有研究不同,我们的方法利用额外说话者提示有效学习从自然语言描述到多样说话者声学特征的映射。主观评估结果表明,相比不使用说话者提示的方法,所提方法能更精准地控制说话者特征。音频样本详见https://reppy4620.github.io/demo.promptttspp/。