Laughter is one of the most expressive and natural aspects of human speech, conveying emotions, social cues, and humor. However, most text-to-speech (TTS) systems lack the ability to produce realistic and appropriate laughter sounds, limiting their applications and user experience. While there have been prior works to generate natural laughter, they fell short in terms of controlling the timing and variety of the laughter to be generated. In this work, we propose ELaTE, a zero-shot TTS that can generate natural laughing speech of any speaker based on a short audio prompt with precise control of laughter timing and expression. Specifically, ELaTE works on the audio prompt to mimic the voice characteristic, the text prompt to indicate the contents of the generated speech, and the input to control the laughter expression, which can be either the start and end times of laughter, or the additional audio prompt that contains laughter to be mimicked. We develop our model based on the foundation of conditional flow-matching-based zero-shot TTS, and fine-tune it with frame-level representation from a laughter detector as additional conditioning. With a simple scheme to mix small-scale laughter-conditioned data with large-scale pre-training data, we demonstrate that a pre-trained zero-shot TTS model can be readily fine-tuned to generate natural laughter with precise controllability, without losing any quality of the pre-trained zero-shot TTS model. Through objective and subjective evaluations, we show that ELaTE can generate laughing speech with significantly higher quality and controllability compared to conventional models. See https://aka.ms/elate/ for demo samples.
翻译:笑声是人类语音中最具表现力和自然性的方面之一,能传达情感、社交暗示和幽默。然而,大多数文本转语音(TTS)系统缺乏生成逼真且恰当笑声的能力,这限制了它们的应用和用户体验。尽管已有先前研究致力于生成自然笑声,但在控制笑声生成的时间点和多样性方面仍显不足。在本工作中,我们提出ELaTE,一种零样本TTS系统,它能够基于简短音频提示生成任意说话者的自然笑声语音,并精确控制笑声的时间和表达方式。具体而言,ELaTE处理音频提示以模仿声音特征,处理文本提示以指示生成语音的内容,并处理输入以控制笑声表达——该输入可以是笑声的起始和结束时间,也可以是包含待模仿笑声的额外音频提示。我们基于条件流匹配零样本TTS的基础开发模型,并通过来自笑声检测器的帧级表示作为额外条件进行微调。借助一种简单方案,将小规模笑声条件数据与大规模预训练数据混合,我们证明预训练的零样本TTS模型可以轻松微调,以生成具有精确可控性的自然笑声,同时不损失预训练零样本TTS模型的任何质量。通过客观和主观评估,我们表明ELaTE能够生成相比传统模型显著更高质量和可控性的笑声语音。演示样本请参见 https://aka.ms/elate/。