Transducer is one of the mainstream frameworks for streaming speech recognition. There is a performance gap between the streaming and non-streaming transducer models due to limited context. To reduce this gap, an effective way is to ensure that their hidden and output distributions are consistent, which can be achieved by hierarchical knowledge distillation. However, it is difficult to ensure the distribution consistency simultaneously because the learning of the output distribution depends on the hidden one. In this paper, we propose an adaptive two-stage knowledge distillation method consisting of hidden layer learning and output layer learning. In the former stage, we learn hidden representation with full context by applying mean square error loss function. In the latter stage, we design a power transformation based adaptive smoothness method to learn stable output distribution. It achieved 19\% relative reduction in word error rate, and a faster response for the first token compared with the original streaming model in LibriSpeech corpus.
翻译:Transducer是流式语音识别的主流框架之一。由于上下文受限,流式与非流式Transducer模型之间存在性能差距。为缩小这一差距,确保其隐层和输出分布的一致性是一种有效方法,这可通过层次化知识蒸馏实现。然而,由于输出分布的学习依赖于隐层分布,同时保证两者的一致性较为困难。本文提出一种自适应两阶段知识蒸馏方法,包含隐层学习和输出层学习两个阶段。在第一阶段,我们采用均方误差损失函数学习包含完整上下文的隐层表示。在第二阶段,我们设计了一种基于幂变换的自适应平滑方法,以学习稳定的输出分布。在LibriSpeech语料库上,该方法相较原始流式模型实现了19%的词错误率相对降低,并加快了首个令牌的响应速度。