End-to-end (e2e) systems have recently gained wide popularity in automatic speech recognition. However, these systems do generally not provide well-calibrated word-level confidences. In this paper, we propose Hystoc, a simple method for obtaining word-level confidences from hypothesis-level scores. Hystoc is an iterative alignment procedure which turns hypotheses from an n-best output of the ASR system into a confusion network. Eventually, word-level confidences are obtained as posterior probabilities in the individual bins of the confusion network. We show that Hystoc provides confidences that correlate well with the accuracy of the ASR hypothesis. Furthermore, we show that utilizing Hystoc in fusion of multiple e2e ASR systems increases the gains from the fusion by up to 1\,\% WER absolute on Spanish RTVE2020 dataset. Finally, we experiment with using Hystoc for direct fusion of n-best outputs from multiple systems, but we only achieve minor gains when fusing very similar systems.
翻译:端到端(e2e)系统近年来在自动语音识别领域广受欢迎。然而,这些系统通常无法提供校准良好的词级置信度。本文提出Hystoc,一种从假设级得分中获取词级置信度的简单方法。Hystoc通过迭代对齐过程,将ASR系统n-best输出中的假设转换为混淆网络。最终,词级置信度作为混淆网络各个区间中的后验概率获得。实验表明,Hystoc提供的置信度与ASR假设的准确性具有良好的相关性。此外,在多e2e ASR系统的融合中应用Hystoc,可在西班牙语RTVE2020数据集上将融合增益提升至多达1%的绝对词错误率(WER)。最后,我们尝试利用Hystoc直接融合多个系统的n-best输出,但在融合非常相似的系统时仅获得微小增益。