Natural language processing (NLP) and speech technologies have made significant progress in recent years; however, they remain largely focused on standardized language varieties. Dialects, despite their cultural significance and widespread use, are underrepresented in linguistic resources and computational models, resulting in performance disparities. To address this gap, we introduce Saar-Voice, a six-hour speech corpus for the Saarbrücken dialect of German. The dataset was created by first collecting text through digitized books and locally sourced materials. A subset of this text was recorded by nine speakers, and we conducted analyses on both the textual and speech components to assess the dataset's characteristics and quality. We discuss methodological challenges related to orthographic and speaker variation, and explore grapheme-to-phoneme (G2P) conversion. The resulting corpus provides aligned textual and audio representations. This serves as a foundation for future research on dialect-aware text-to-speech (TTS), particularly in low-resource scenarios, including zero-shot and few-shot model adaptation.
翻译:自然语言处理与语音技术近年来取得了显著进展,但仍主要聚焦于标准语言变体。方言尽管具有重要文化意义且使用广泛,却在语言资源和计算模型中的代表性不足,导致性能差异。为弥补这一缺口,我们提出了Saar-Voice——一个面向德语萨尔布吕肯方言的六小时语音语料库。该数据集首先通过数字化书籍和本地来源材料收集文本,并由九位说话人录制了其中一部分文本。我们对文本与语音组分进行了分析以评估数据集的特征与质量。讨论了与正字法及说话人变体相关的方法论挑战,并探索了字素-音素转换。最终生成的语料库提供了对齐的文本与音频表征,为未来方言感知的文本转语音研究(尤其在低资源场景中,含零样本和少样本模型适配)奠定了基础。