Several high-resource Text to Speech (TTS) systems currently produce natural, well-established human-like speech. In contrast, low-resource languages, including Arabic, have very limited TTS systems due to the lack of resources. We propose a fully unsupervised method for building TTS, including automatic data selection and pre-training/fine-tuning strategies for TTS training, using broadcast news as a case study. We show how careful selection of data, yet smaller amounts, can improve the efficiency of TTS system in generating more natural speech than a system trained on a bigger dataset. We adopt to propose different approaches for the: 1) data: we applied automatic annotations using DNSMOS, automatic vowelization, and automatic speech recognition (ASR) for fixing transcriptions' errors; 2) model: we used transfer learning from high-resource language in TTS model and fine-tuned it with one hour broadcast recording then we used this model to guide a FastSpeech2-based Conformer model for duration. Our objective evaluation shows 3.9% character error rate (CER), while the groundtruth has 1.3% CER. As for the subjective evaluation, where 1 is bad and 5 is excellent, our FastSpeech2-based Conformer model achieved a mean opinion score (MOS) of 4.4 for intelligibility and 4.2 for naturalness, where many annotators recognized the voice of the broadcaster, which proves the effectiveness of our proposed unsupervised method.
翻译:当前多个高资源文本转语音(TTS)系统已能生成自然流畅、类人化的语音。然而,包括阿拉伯语在内的低资源语言因资源匮乏,其TTS系统建设严重受限。本文以广播新闻为案例,提出了一种完全无监督的TTS构建方法,涵盖自动数据选择及针对TTS训练的预训练/微调策略。我们证明:相较于使用更大数据集训练的系统,通过精心选择较小规模的数据可有效提升TTS系统生成更自然语音的效能。我们针对以下环节提出不同方案:1)数据层面:采用DNSMOS自动标注、自动注音及自动语音识别(ASR)修正转录错误;2)模型层面:通过高资源语言的TTS迁移学习,利用一小时广播录音进行微调,并以此模型指导基于FastSpeech2的Conformer模型实现时长控制。客观评测显示字符错误率(CER)为3.9%(基准真值1.3%)。主观评测(1-5分制)中,我们提出的FastSpeech2-Conformer模型在可懂度与自然度上分别获得4.4分和4.2分的平均意见得分(MOS),多名标注者能识别出广播员原声,充分验证了所提无监督方法的有效性。