This paper proposes a novel automatic speech recognition (ASR) system that can transcribe individual speaker's speech while identifying whether they are target or non-target speakers from multi-talker overlapped speech. Target-speaker ASR systems are a promising way to only transcribe a target speaker's speech by enrolling the target speaker's information. However, in conversational ASR applications, transcribing both the target speaker's speech and non-target speakers' ones is often required to understand interactive information. To naturally consider both target and non-target speakers in a single ASR model, our idea is to extend autoregressive modeling-based multi-talker ASR systems to utilize the enrollment speech of the target speaker. Our proposed ASR is performed by recursively generating both textual tokens and tokens that represent target or non-target speakers. Our experiments demonstrate the effectiveness of our proposed method.
翻译:本文提出了一种新颖的自动语音识别(ASR)系统,该系统能够在多人重叠语音中转录单个说话人的语音,同时识别该说话人是目标说话人还是非目标说话人。目标说话人ASR系统通过输入目标说话人的语音信息来仅转录目标说话人的语音,是一种有前景的方法。然而,在对话ASR应用中,通常需要同时转录目标说话人和非目标说话人的语音,以理解交互信息。为了在单一ASR模型中自然兼顾目标与非目标说话人,我们的思路是扩展基于自回归建模的多说话人ASR系统,使其能够利用目标说话人的注册语音。所提出的ASR系统通过递归生成文本标记以及表示目标或非目标说话人的标记来执行。实验证明了我们提出方法的有效性。