With the success of the first Multi-channel Multi-party Meeting Transcription challenge (M2MeT), the second M2MeT challenge (M2MeT 2.0) held in ASRU2023 particularly aims to tackle the complex task of speaker-attributed ASR (SA-ASR), which directly addresses the practical and challenging problem of "who spoke what at when" at typical meeting scenario. We particularly established two sub-tracks. 1) The fixed training condition sub-track, where the training data is constrained to predetermined datasets, but participants can use any open-source pre-trained model. 2) The open training condition sub-track, which allows for the use of all available data and models. In addition, we release a new 10-hour test set for challenge ranking. This paper provides an overview of the dataset, track settings, results, and analysis of submitted systems, as a benchmark to show the current state of speaker-attributed ASR.
翻译:随着首届多通道多方会议转录挑战赛(M2MeT)的成功举办,第二届M2MeT挑战赛(M2MeT 2.0)在ASRU2023上特别旨在解决说话人归属语音识别(SA-ASR)这一复杂任务,该任务直接针对典型会议场景中"谁在何时说了什么"这一实际且具有挑战性的问题。我们特别设立了以下两个子赛道:1) 固定训练条件赛道,其中训练数据被限制在预定数据集内,但参与者可使用任意开源预训练模型;2) 开放训练条件赛道,该赛道允许使用所有可用数据和模型。此外,我们发布了一个全新的10小时测试集用于挑战排名。本文概述了数据集、赛道设置、结果及提交系统的分析,以此作为展示说话人归属语音识别当前水平的基准。