In this paper, alternating weak triphone/BPE alignment supervision is proposed to improve end-to-end model training. Towards this end, triphone and BPE alignments are extracted using a pre-existing hybrid ASR system. Then, regularization effect is obtained by cross-entropy based intermediate auxiliary losses computed on such alignments at a mid-layer representation of the encoder for triphone alignments and at the encoder for BPE alignments. Weak supervision is achieved through strong label smoothing with parameter of 0.5. Experimental results on TED-LIUM 2 indicate that either triphone or BPE alignment based weak supervision improves ASR performance over standard CTC auxiliary loss. Moreover, their combination lowers the word error rate further. We also investigate the alternation of the two auxiliary tasks during model training, and additional performance gain is observed. Overall, the proposed techniques result in over 10% relative error rate reduction over a CTC-regularized baseline system.
翻译:本文提出一种基于交替弱三音素/BPE对齐监督的方法,以改进端到端模型训练。为此,首先利用预训练的混合语音识别系统提取三音素和BPE对齐信息。随后,通过交叉熵损失在编码器中间层(针对三音素对齐)和编码器输出层(针对BPE对齐)计算中间辅助损失,从而获得正则化效果。弱监督通过参数为0.5的强标签平滑实现。在TED-LIUM 2数据集上的实验表明,基于三音素或BPE对齐的弱监督方法均优于标准CTC辅助损失。进一步地,两种对齐方式的组合可进一步降低词错误率。我们还研究了模型训练中两种辅助任务的交替策略,观察到额外性能提升。总体而言,所提方法相比CTC正则化基线系统实现了超过10%的相对错误率降低。