In conversational speech separation and recognition tasks, close-talk microphones are typically attached to each speaker during training data collection to capture near-field, close-talk mixture signals, in addition to using far-field microphones to record far-field mixture signals. Each such close-talk mixture exhibits a reasonably high energy level for the wearer and could intuitively serve as weak supervision for training far-field speech separation models directly on real-recorded far-field signals. However, they are not sufficiently clean for this purpose, as they often contain strong cross-talk speech from other speakers in addition to background noise. To address this, we propose cross-talk reduction (CTR), a task aiming to isolate the wearer's speech from each close-talk mixture, and a novel method called CTRnet, which can be trained directly on real-recorded pairs of close-talk and far-field mixtures to accomplish CTR. Building on CTRnet, we further propose pseudo-label based far-field speech separation (PuLSS), which uses CTRnet's estimated clean speech as pseudo-labels to train models for separating far-field mixtures. A key advantage of the proposed framework is that both CTRnet and PuLSS can be trained on real-recorded data from the target domain, addressing the generalization gap commonly observed when models are trained exclusively on simulated data. On the CHiME-6 dataset, our framework achieves state-of-the-art ASR performance under both oracle and estimated speaker diarization, surpassing all CHiME-{7,8} challenge submissions. To our knowledge, it is the first neural speech separation method that substantially outperforms guided source separation on real conversational "speech-in-the-wild" data.
翻译:摘要:在对话语音分离与识别任务中,训练数据采集时通常会在每位说话人身上佩戴近讲麦克风以捕获近场混合信号,同时使用远场麦克风记录远场混合信号。每个近讲混合信号中佩戴者的语音能量水平相对较高,因此理论上可直接作为弱监督信号,用于训练基于真实远场信号的语音分离模型。然而,这些近讲信号中常包含来自其他说话人的强烈交叉干扰语音以及背景噪声,因而无法直接用于此目的。为解决该问题,我们提出交叉干扰消除(CTR)任务——旨在从每个近讲混合信号中分离出佩戴者语音,并设计了一种名为CTRnet的新方法,该方法可直接利用真实录制的近讲-远场混合信号对进行训练以实现CTR。基于CTRnet,我们进一步提出基于伪标签的远场语音分离方法(PuLSS),该方法将CTRnet估计的纯净语音作为伪标签,用于训练远场混合信号分离模型。本框架的关键优势在于:CTRnet和PuLSS均可直接使用目标域的真实录制数据进行训练,从而有效弥补传统仅依赖仿真数据训练模型时常见的泛化差距。在CHiME-6数据集上,本框架在先知说话人分割与估计说话人分割两种场景下均取得了最优的ASR性能,超越了CHiME-{7,8}挑战赛的所有提交方案。据我们所知,这是首个在真实“野外”对话数据上显著优于引导源分离方法的神经语音分离方法。