Speech technologies rely on capturing a speaker's voice variability while obtaining comprehensive language information. Textual prompts and sentence selection methods have been proposed in the literature to comprise such adequate phonetic data, referred to as a phonetically rich \textit{corpus}. However, they are still insufficient for acoustic modeling, especially critical for languages with limited resources. Hence, this paper proposes a novel approach and outlines the methodological aspects required to create a \textit{corpus} with broad phonetic coverage for a low-resourced language, Brazilian Portuguese. Our methodology includes text dataset collection up to a sentence selection algorithm based on triphone distribution. Furthermore, we propose a new phonemic classification according to acoustic-articulatory speech features since the absolute number of distinct triphones, or low-probability triphones, does not guarantee an adequate representation of every possible combination. Using our algorithm, we achieve a 55.8\% higher percentage of distinct triphones -- for samples of similar size -- while the currently available phonetic-rich corpus, CETUC and TTS-Portuguese, 12.6\% and 12.3\% in comparison to a non-phonetically rich dataset.
翻译:语音技术依赖于在获取全面语言信息的同时捕获说话者的语音变异性。文献中已提出文本提示和句子选择方法,用于构成此类充分的语音数据,即所谓的语音丰富语料库。然而,这些方法对于声学建模仍显不足,尤其对于资源受限的语言至关重要。因此,本文提出一种新方法,并概述了为低资源语言——巴西葡萄牙语——创建具有广泛语音覆盖范围的语料库所需的方法论要点。我们的方法论包括从文本数据集收集到基于三音素分布的句子选择算法。此外,我们根据声学-发音语音特征提出了一种新的音位分类,因为不同三音素的绝对数量或低概率三音素并不能保证每种可能组合的充分表示。使用我们的算法,在相似规模样本下,我们实现了比现有语音丰富语料库CETUC和TTS-Portuguese分别高出55.8%的不同三音素比例——后者相较于非语音丰富数据集的比例分别为12.6%和12.3%。