This paper addresses the problem of improving POS tagging of transcripts of speech from clinical populations. In contrast to prior work on parsing and POS tagging of transcribed speech, we do not make use of an in domain treebank for training. Instead, we train on an out of domain treebank of newswire using data augmentation techniques to make these structures resemble natural, spontaneous speech. We trained a parser with and without the augmented data and tested its performance using manually validated POS tags in clinical speech produced by patients with various types of neurodegenerative conditions.
翻译:本文针对临床人群语音转录文本的词性标注改进问题展开研究。与以往基于口语转录文本的句法分析和词性标注工作不同,我们并未使用领域内树库进行训练,而是采用数据增强技术对新闻语料库(领域外树库)进行处理,使其结构接近自然口语。我们分别使用增强数据与原始数据训练解析器,并以不同神经退行性疾病患者临床语音的手动验证词性标注为基准,评估模型性能。