Style transfer TTS has shown impressive performance in recent years. However, style control is often restricted to systems built on expressive speech recordings with discrete style categories. In practical situations, users may be interested in transferring style by typing text descriptions of desired styles, without the reference speech in the target style. The text-guided content generation techniques have drawn wide attention recently. In this work, we explore the possibility of controllable style transfer with natural language descriptions. To this end, we propose PromptStyle, a text prompt-guided cross-speaker style transfer system. Specifically, PromptStyle consists of an improved VITS and a cross-modal style encoder. The cross-modal style encoder constructs a shared space of stylistic and semantic representation through a two-stage training process. Experiments show that PromptStyle can achieve proper style transfer with text prompts while maintaining relatively high stability and speaker similarity. Audio samples are available in our demo page.
翻译:摘要:近年来,风格迁移文本到语音技术展现出令人瞩目的性能。然而,风格控制通常局限于基于具有离散风格类别的情感语音录音构建的系统。在实际场景中,用户可能更倾向于通过输入期望风格的文本描述来实现风格迁移,而无需提供目标风格的参考语音。文本引导的内容生成技术近期引起了广泛关注。本研究探索了基于自然语言描述进行可控风格迁移的可能性。为此,我们提出PromptStyle——一种文本提示引导的跨说话人风格迁移系统。具体而言,PromptStyle由改进版VITS和跨模态风格编码器构成。该跨模态风格编码器通过两阶段训练过程构建风格表征与语义表征的共享空间。实验表明,PromptStyle能够在保持较高稳定性和说话人相似度的同时,通过文本提示实现合理的风格迁移。音频样本可访问我们的演示页面获取。