Previous Multimodal Information based Speech Processing (MISP) challenges mainly focused on audio-visual speech recognition (AVSR) with commendable success. However, the most advanced back-end recognition systems often hit performance limits due to the complex acoustic environments. This has prompted a shift in focus towards the Audio-Visual Target Speaker Extraction (AVTSE) task for the MISP 2023 challenge in ICASSP 2024 Signal Processing Grand Challenges. Unlike existing audio-visual speech enhance-ment challenges primarily focused on simulation data, the MISP 2023 challenge uniquely explores how front-end speech processing, combined with visual clues, impacts back-end tasks in real-world scenarios. This pioneering effort aims to set the first benchmark for the AVTSE task, offering fresh insights into enhancing the ac-curacy of back-end speech recognition systems through AVTSE in challenging and real acoustic environments. This paper delivers a thorough overview of the task setting, dataset, and baseline system of the MISP 2023 challenge. It also includes an in-depth analysis of the challenges participants may encounter. The experimental results highlight the demanding nature of this task, and we look forward to the innovative solutions participants will bring forward.
翻译:以往的多模态信息语音处理(MISP)挑战赛主要聚焦于音视频语音识别(AVSR),并取得了显著成效。然而,由于复杂声学环境的限制,最先进的后端识别系统常常遭遇性能瓶颈。这促使ICASSP 2024信号处理大挑战赛中的MISP 2023挑战赛将重心转向音视频目标说话人提取(AVTSE)任务。与现有主要基于仿真数据的音视频语音增强挑战不同,MISP 2023挑战赛独树一帜地探索了前端语音处理结合视觉线索在真实场景中对后端任务的影响。这项开创性工作旨在为AVTSE任务建立首个基准,通过AVTSE在复杂真实声学环境中提升后端语音识别系统的精度,提供全新见解。本文全面概述了MISP 2023挑战赛的任务设定、数据集与基线系统,并深入分析了参与者可能面临的挑战。实验结果表明了该任务的艰巨性,我们期待参与者带来创新解决方案。