Punctuation restoration is an important task in automatic speech recognition (ASR) which aim to restore the syntactic structure of generated ASR texts to improve readability. While punctuated texts are abundant from written documents, the discrepancy between written punctuated texts and ASR texts limits the usability of written texts in training punctuation restoration systems for ASR texts. This paper proposes a reinforcement learning method to exploit in-topic written texts and recent advances in large pre-trained generative language models to bridge this gap. The experiments show that our method achieves state-of-the-art performance on the ASR test set on two benchmark datasets for punctuation restoration.
翻译:标点恢复是自动语音识别(ASR)中的一项重要任务,旨在恢复生成的ASR文本的句法结构,以提升文本可读性。尽管书面文档中存在大量带标点的文本,但书面带标点文本与ASR文本之间的差异限制了在训练ASR文本标点恢复系统时对书面文本的可用性。本文提出一种强化学习方法,利用主题内书面文本及大型预训练生成语言模型的最新进展来弥合这一差距。实验表明,我们的方法在两个标点恢复基准数据集的ASR测试集上达到了最先进的性能。