Context cues carry information which can improve multi-turn interactions in automatic speech recognition (ASR) systems. In this paper, we introduce a novel mechanism inspired by hyper-prompting to fuse textual context with acoustic representations in the attention mechanism. Results on a test set with multi-turn interactions show that our method achieves 5.9% relative word error rate reduction (rWERR) over a strong baseline. We show that our method does not degrade in the absence of context and leads to improvements even if the model is trained without context. We further show that leveraging a pre-trained sentence-piece model for context embedding generation can outperform an external BERT model.
翻译:上下文线索携带的信息可提升自动语音识别(ASR)系统中多轮交互的性能。本文提出一种受超提示启发的创新机制,将文本上下文与注意力机制中的声学表征相融合。在多轮交互测试集上的实验结果表明,与强基线模型相比,我们的方法实现了5.9%的相对词错误率降低(rWERR)。研究显示,该方法在无上下文场景下性能不会退化,且即便模型未经上下文训练也能带来改进。进一步实验证明,利用预训练句子片段模型生成上下文嵌入的效果优于外部BERT模型。