Text Normalization is an integral part of any text-to-speech synthesis system. In a natural language text, there are elements such as numbers, dates, abbreviations, etc. that belong to other semiotic classes. They are called non-standard words (NSW) and need to be expanded into ordinary words. For this purpose, it is necessary to identify the semiotic class of each NSW. The taxonomy of semiotic classes adapted to the Lithuanian language is presented in the work. Sets of rules are created for detecting and expanding NSWs based on regular expressions. Experiments with three completely different data sets were performed and the accuracy was assessed. Causes of errors are explained and recommendations are given for the development of text normalization rules.
翻译:文本规范化是任何文本转语音合成系统中不可或缺的组成部分。在自然语言文本中,存在数字、日期、缩写等属于其他符号类别的元素。这些元素被称为非标准词(NSW),需要将其扩展为普通词汇。为此,必须识别每个非标准词的符号类别。本文提出了适用于立陶宛语的符号类别分类体系,并基于正则表达式构建了用于检测和扩展非标准词的规则集。通过对三个完全不同的数据集进行实验,评估了方法的准确性。同时分析了错误产生的原因,并为制定文本规范化规则提出了建议。