At present, Text-to-speech (TTS) systems that are trained with high-quality transcribed speech data using end-to-end neural models can generate speech that is intelligible, natural, and closely resembles human speech. These models are trained with relatively large single-speaker professionally recorded audio, typically extracted from audiobooks. Meanwhile, due to the scarcity of freely available speech corpora of this kind, a larger gap exists in Arabic TTS research and development. Most of the existing freely available Arabic speech corpora are not suitable for TTS training as they contain multi-speaker casual speech with variations in recording conditions and quality, whereas the corpus curated for speech synthesis are generally small in size and not suitable for training state-of-the-art end-to-end models. In a move towards filling this gap in resources, we present a speech corpus for Classical Arabic Text-to-Speech (ClArTTS) to support the development of end-to-end TTS systems for Arabic. The speech is extracted from a LibriVox audiobook, which is then processed, segmented, and manually transcribed and annotated. The final ClArTTS corpus contains about 12 hours of speech from a single male speaker sampled at 40100 kHz. In this paper, we describe the process of corpus creation and provide details of corpus statistics and a comparison with existing resources. Furthermore, we develop two TTS systems based on Grad-TTS and Glow-TTS and illustrate the performance of the resulting systems via subjective and objective evaluations. The corpus will be made publicly available at www.clartts.com for research purposes, along with the baseline TTS systems demo.
翻译:目前,利用高质量转录语音数据训练的端到端神经模型所生成的文本转语音(TTS)系统,能够产生清晰、自然且高度接近人类语音的合成语音。这类模型通常使用由专业录音制作的大型单说话人音频(多从有声书中提取)进行训练。然而,由于此类可免费获取的语音语料库匮乏,阿拉伯语TTS研究与开发面临较大差距。现有大多数可免费获取的阿拉伯语语音语料库因包含多说话人随意语音且录音条件与质量参差不齐,不适用于TTS训练,而专为语音合成整理的语料库规模普遍较小,不足以训练先进的端到端模型。为填补这一资源缺口,我们发布了面向古典阿拉伯语文本转语音(ClArTTS)的语音语料库,以支持阿拉伯语端到端TTS系统的开发。语音数据提取自LibriVox有声书,经处理、切分、人工转录与标注后,最终形成ClArTTS语料库,包含约12小时的单人男性说话人语音,采样率为40100 kHz。本文描述了语料库的创建过程,提供了语料库统计详情及与现有资源的比较。此外,我们基于Grad-TTS和Glow-TTS开发了两套TTS系统,并通过主观与客观评估展示了系统性能。该语料库将连同基线TTS系统演示在www.clartts.com上公开供研究使用。