Achieving high accuracy with low latency has always been a challenge in streaming end-to-end automatic speech recognition (ASR) systems. By attending to more future contexts, a streaming ASR model achieves higher accuracy but results in larger latency, which hurts the streaming performance. In the Mask-CTC framework, an encoder network is trained to learn the feature representation that anticipates long-term contexts, which is desirable for streaming ASR. Mask-CTC-based encoder pre-training has been shown beneficial in achieving low latency and high accuracy for triggered attention-based ASR. However, the effectiveness of this method has not been demonstrated for various model architectures, nor has it been verified that the encoder has the expected look-ahead capability to reduce latency. This study, therefore, examines the effectiveness of Mask-CTCbased pre-training for models with different architectures, such as Transformer-Transducer and contextual block streaming ASR. We also discuss the effect of the proposed pre-training method on obtaining accurate output spike timing.
翻译:在流式端到端自动语音识别(ASR)系统中,实现高精度与低延迟始终是一项挑战。通过关注更多未来上下文,流式ASR模型能够获得更高的准确率,但同时也导致更大的延迟,从而影响流式性能。在掩码CTC框架中,编码器网络被训练以学习能够预测长期上下文的特征表示,这对于流式ASR而言具有理想特性。基于掩码CTC的编码器预训练已被证明有益于基于触发注意力的ASR实现低延迟和高准确率。然而,该方法对不同模型架构的有效性尚未得到证实,也未验证编码器是否具备预期的前瞻能力以减少延迟。因此,本研究考察了掩码CTC预训练对不同架构模型(如Transformer-Transducer和上下文块流式ASR)的有效性。同时,我们还讨论了所提出的预训练方法对获取精确输出脉冲时序的影响。