Persian remains substantially underrepresented in open speech-text resources, limiting progress in multi-speaker text-to-speech (TTS), speech-language modelling, and low-resource speech processing. We introduce ParsVoice, the largest publicly available Persian speech-text corpus tailored for training multi-speaker TTS systems, along with a scalable pipeline to construct high-quality speech-text data from long-form audiobook recordings. The pipeline combines a fine-tuned ParsBERT sentence-completion classifier, ASR-based boundary optimization, punctuation restoration, speaker identification, and a multi-dimensional quality assessment that covers both audio and Persian-specific text properties. The resulting release contains a 2,200-hour TTS-ready subset with 1.36 million aligned segments from 1,815 automatically identified speaker IDs, making it more than 25 times larger than the previously largest open Persian TTS dataset. To validate the corpus, we fine-tune XTTS, a zero-shot multilingual TTS model that operates directly on raw Persian text without phoneme representations, achieving a naturalness MOS of 3.6/5 and speaker similarity MOS of 4.0/5. The ParsVoice dataset is publicly available at: https://huggingface.co/datasets/MohammadJRanjbar/ParsVoice.
翻译:波斯语在公开语音-文本资源中仍严重不足,限制了多说话人文本转语音(TTS)、语音语言建模及低资源语音处理领域的发展。我们提出ParsVoice——当前最大的公开波斯语语音-文本语料库,专为训练多说话人TTS系统设计,并配套开发了可扩展流水线,用于从长篇有声书录音中构建高质量语音-文本数据。该流水线融合了经过微调的ParsBERT句子补全分类器、基于ASR的边界优化、标点符号恢复、说话人识别,以及涵盖音频与波斯语特定文本属性的多维质量评估。最终发布的语料包含2200小时的TTS就绪子集,含136万个对齐片段及1815个自动识别说话人ID,规模较此前最大公开波斯语TTS数据集扩大逾25倍。为验证语料质量,我们对XTTS(直接处理原始波斯语文本、无需音素表征的零样本多语言TTS模型)进行微调,自然度MOS得分为3.6/5,说话人相似度MOS得分为4.0/5。ParsVoice数据集公开访问地址:https://huggingface.co/datasets/MohammadJRanjbar/ParsVoice