End-to-end multi-talker speech recognition has garnered great interest as an effective approach to directly transcribe overlapped speech from multiple speakers. Current methods typically adopt either 1) single-input multiple-output (SIMO) models with a branched encoder, or 2) single-input single-output (SISO) models based on attention-based encoder-decoder architecture with serialized output training (SOT). In this work, we propose a Cross-Speaker Encoding (CSE) network to address the limitations of SIMO models by aggregating cross-speaker representations. Furthermore, the CSE model is integrated with SOT to leverage both the advantages of SIMO and SISO while mitigating their drawbacks. To the best of our knowledge, this work represents an early effort to integrate SIMO and SISO for multi-talker speech recognition. Experiments on the two-speaker LibrispeechMix dataset show that the CES model reduces word error rate (WER) by 8% over the SIMO baseline. The CSE-SOT model reduces WER by 10% overall and by 16% on high-overlap speech compared to the SOT model.
翻译:端到端多说话人语音识别作为直接转录多个说话人重叠语音的有效方法,已引起广泛关注。现有方法通常采用两种范式:1)带有分支编码器的单输入多输出(SIMO)模型,或2)基于注意力编码器-解码器架构并采用序列化输出训练(SOT)的单输入单输出(SISO)模型。本文提出跨说话人编码(CSE)网络,通过聚合跨说话人表征来克服SIMO模型的局限性。进一步地,我们将CSE模型与SOT相结合,在融合SIMO与SISO优势的同时缓解其各自缺陷。据我们所知,这是首个探索SIMO与SISO融合的多说话人语音识别研究。在双说话人LibrispeechMix数据集上的实验表明,CSE模型相比SIMO基线模型词错误率(WER)降低8%;CSE-SOT模型相较SOT模型总体WER降低10%,在高重叠语音片段上WER降低16%。