With read-aloud speech synthesis achieving high naturalness scores, there is a growing research interest in synthesising spontaneous speech. However, human spontaneous face-to-face conversation has both spoken and non-verbal aspects (here, co-speech gestures). Only recently has research begun to explore the benefits of jointly synthesising these two modalities in a single system. The previous state of the art used non-probabilistic methods, which fail to capture the variability of human speech and motion, and risk producing oversmoothing artefacts and sub-optimal synthesis quality. We present the first diffusion-based probabilistic model, called Diff-TTSG, that jointly learns to synthesise speech and gestures together. Our method can be trained on small datasets from scratch. Furthermore, we describe a set of careful uni- and multi-modal subjective tests for evaluating integrated speech and gesture synthesis systems, and use them to validate our proposed approach. Please see https://shivammehta25.github.io/Diff-TTSG/ for video examples, data, and code.
翻译:随着朗读式语音合成达到较高的自然度评分,自发语音合成领域的研究兴趣日益增长。然而,人类自发面对面交流兼具口语与非语言方面(此处指伴随语音的手势)。直到近期,研究才开始探索在单一系统中联合合成这两种模态的益处。先前的最优方法采用非概率模型,难以捕捉人类语音与动作的变异性,且易产生过度平滑伪影及次优合成质量。我们首次提出基于扩散的概率模型Diff-TTSG,该模型可联合学习合成语音与手势。我们的方法能够从头开始在小数据集上训练。此外,我们描述了一套用于评估集成语音与手势合成系统的谨慎的单模态与多模态主观测试,并以此验证所提方法。视频示例、数据及代码详见https://shivammehta25.github.io/Diff-TTSG/。