Speaker-attributed automatic speech recognition (SA-ASR) improves the accuracy and applicability of multi-speaker ASR systems in real-world scenarios by assigning speaker labels to transcribed texts. However, SA-ASR poses unique challenges due to factors such as speaker overlap, speaker variability, background noise, and reverberation. In this study, we propose PP-MeT system, a real-world personalized prompt based meeting transcription system, which consists of a clustering system, target-speaker voice activity detection (TS-VAD), and TS-ASR. Specifically, we utilize target-speaker embedding as a prompt in TS-VAD and TS-ASR modules in our proposed system. In constrast with previous system, we fully leverage pre-trained models for system initialization, thereby bestowing our approach with heightened generalizability and precision. Experiments on M2MeT2.0 Challenge dataset show that our system achieves a cp-CER of 11.27% on the test set, ranking first in both fixed and open training conditions.
翻译:说话人属性自动语音识别(SA-ASR)通过为转录文本分配说话人标签,提升了多说话人ASR系统在真实场景中的准确性和适用性。然而,SA-ASR由于说话人重叠、说话人差异、背景噪声及混响等因素面临着独特挑战。本研究提出PP-MeT系统,这是一个基于真实世界个性化提示的会议转录系统,包含聚类系统、目标说话人语音活动检测(TS-VAD)和TS-ASR模块。具体而言,我们在所提系统的TS-VAD和TS-ASR模块中利用目标说话人嵌入作为提示。与以往系统相比,我们充分利用预训练模型进行系统初始化,从而赋予方法更高的泛化能力和精度。在M2MeT2.0挑战数据集上的实验表明,我们的系统在测试集上达到11.27%的cp-CER,在固定和开放训练条件下均排名第一。