In multilingual societies, social conversations often involve code-mixed speech. The current speech technology may not be well equipped to extract information from multi-lingual multi-speaker conversations. The DISPLACE challenge entails a first-of-kind task to benchmark speaker and language diarization on the same data, as the data contains multi-speaker conversations in multilingual code-mixed speech. The challenge attempts to highlight outstanding issues in speaker diarization (SD) in multilingual settings with code-mixing. Further, language diarization (LD) in multi-speaker settings also introduces new challenges, where the system has to disambiguate speaker switches with code switches. For this challenge, a natural multilingual, multi-speaker conversational dataset is distributed for development and evaluation purposes. The systems are evaluated on single-channel far-field recordings. We also release a baseline system and report the highlights of the system submissions.
翻译:在多语种社会中,社交对话常涉及语码混合的语音。当前语音技术可能难以从多语种、多说话人对话中有效提取信息。DISPLACE挑战首次提出在同一数据上对说话人和语言分离进行基准测试的任务,该数据包含多语种语码混合的多说话人对话。该挑战旨在突出多语种环境下语码混合情形中的说话人分离(SD)问题。此外,多说话人环境下的语言分离(LD)也带来了新挑战,系统需同时解决说话人切换与语码转换的歧义问题。为支持本挑战,我们发布了一个自然的多语种、多说话人对话数据集用于开发和评估。系统基于单通道远场录音进行评价。我们还发布了基线系统,并概述了各参赛系统的亮点。