Self-supervised learning (SSL) speech representations learned from large amounts of diverse, mixed-quality speech data without transcriptions are gaining ground in many speech technology applications. Prior work has shown that SSL is an effective intermediate representation in two-stage text-to-speech (TTS) for both read and spontaneous speech. However, it is still not clear which SSL and which layer from each SSL model is most suited for spontaneous TTS. We address this shortcoming by extending the scope of comparison for SSL in spontaneous TTS to 6 different SSLs and 3 layers within each SSL. Furthermore, SSL has also shown potential in predicting the mean opinion scores (MOS) of synthesized speech, but this has only been done in read-speech MOS prediction. We extend an SSL-based MOS prediction framework previously developed for scoring read speech synthesis and evaluate its performance on synthesized spontaneous speech. All experiments are conducted twice on two different spontaneous corpora in order to find generalizable trends. Overall, we present comprehensive experimental results on the use of SSL in spontaneous TTS and MOS prediction to further quantify and understand how SSL can be used in spontaneous TTS. Audios samples: https://www.speech.kth.se/tts-demos/sp_ssl_tts
翻译:自监督学习(SSL)语音表征通过大量多样、混合质量的语音数据(无需转录文本)学习得到,正逐步在众多语音技术应用中占据主导地位。先前研究表明,在面向朗读语音和自发性语音的两阶段文本转语音(TTS)系统中,SSL是一种有效的中间表征。然而,目前仍不明确哪种SSL模型及其哪一层最适合自发性TTS。为解决这一不足,我们将自发性TTS中SSL的比较范围扩展至6种不同SSL模型,并分别考察每种模型的3个层级。此外,SSL在预测合成语音的平均意见得分(MOS)方面也展现出潜力,但此方法此前仅应用于朗读语音的MOS预测。我们扩展了先前为评估朗读语音合成而开发的基于SSL的MOS预测框架,并评估其在合成自发性语音上的性能。所有实验均在两个不同的自发性语料库上各进行两次,以寻找可泛化的趋势。整体上,我们展示了SSL在自发性TTS和MOS预测中的全面实验结果,进一步量化和理解SSL如何在自发性TTS中被应用。音频样例:https://www.speech.kth.se/tts-demos/sp_ssl_tts