Despite advancements of end-to-end (E2E) models in speech recognition, named entity recognition (NER) is still challenging but critical for semantic understanding. Previous studies mainly focus on various rule-based or attention-based contextual biasing algorithms. However, their performance might be sensitive to the biasing weight or degraded by excessive attention to the named entity list, along with a risk of false triggering. Inspired by the success of the class-based language model (LM) in NER in conventional hybrid systems and the effective decoupling of acoustic and linguistic information in the factorized neural Transducer (FNT), we propose C-FNT, a novel E2E model that incorporates class-based LMs into FNT. In C-FNT, the LM score of named entities can be associated with the name class instead of its surface form. The experimental results show that our proposed C-FNT significantly reduces error in named entities without hurting performance in general word recognition.
翻译:尽管端到端(E2E)模型在语音识别中取得了进展,但命名实体识别(NER)对于语义理解仍然具有挑战性且至关重要。以往研究主要关注基于规则或注意力的上下文偏置算法,然而其性能可能对偏置权重敏感,或因过度关注命名实体列表而下降,并存在误触发风险。受类语言模型(LM)在传统混合系统NER中的成功应用,以及分解式神经换能器(FNT)有效解耦声学与语言信息的启发,我们提出C-FNT——一种将类语言模型融入FNT的新型E2E模型。在C-FNT中,命名实体的语言模型分数可与名称类别关联而非其表面形式。实验结果表明,所提出的C-FNT在不降低通用词汇识别性能的情况下,显著降低了命名实体的识别错误。