Text normalization - the conversion of text from written to spoken form - is traditionally assumed to be an ill-formed task for language models. In this work, we argue otherwise. We empirically show the capacity of Large-Language Models (LLM) for text normalization in few-shot scenarios. Combining self-consistency reasoning with linguistic-informed prompt engineering, we find LLM based text normalization to achieve error rates around 40\% lower than top normalization systems. Further, upon error analysis, we note key limitations in the conventional design of text normalization tasks. We create a new taxonomy of text normalization errors and apply it to results from GPT-3.5-Turbo and GPT-4.0. Through this new framework, we can identify strengths and weaknesses of GPT-based TN, opening opportunities for future work.
翻译:文本规范化——即将书面文本转换为口语形式的过程——传统上被认为是一项语言模型难以处理的任务。本研究提出相反观点。我们通过实验证明大语言模型在少样本场景下具备文本规范化能力。通过将自一致性推理与基于语言学的提示工程相结合,我们发现基于大语言模型的文本规范化系统错误率比最优系统降低约40%。进一步错误分析显示,传统文本规范化任务设计存在关键局限。我们构建了新的文本规范化错误分类体系,并将其应用于GPT-3.5-Turbo和GPT-4.0的结果分析。通过这一新框架,我们识别出基于GPT的文本规范化的优势与不足,为后续研究开辟了可能性。