Since diarization and source separation of meeting data are closely related tasks, we here propose an approach to perform the two objectives jointly. It builds upon the target-speaker voice activity detection (TS-VAD) diarization approach, which assumes that initial speaker embeddings are available. We replace the final combined speaker activity estimation network of TS-VAD with a network that produces speaker activity estimates at a time-frequency resolution. Those act as masks for source extraction, either via masking or via beamforming. The technique can be applied both for single-channel and multi-channel input and, in both cases, achieves a new state-of-the-art word error rate (WER) on the LibriCSS meeting data recognition task. We further compute speaker-aware and speaker-agnostic WERs to isolate the contribution of diarization errors to the overall WER performance.
翻译:由于会议数据的日志分离与源分离是紧密相关的任务,本文提出了一种联合执行这两项任务的方法。该方法基于目标说话人语音活动检测(TS-VAD)日志分离框架,该框架假设初始说话人嵌入可用。我们将TS-VAD中的最终组合说话人活动估计网络替换为生成时频分辨率下说话人活动估计的网络,这些估计作为源提取的掩码,通过掩码或波束成形实现。该技术可应用于单通道和多通道输入,在两种情况下均在LibriCSS会议数据识别任务上达到了新的词错误率(WER)最优水平。进一步地,我们计算了说话人感知与说话人不可知的WER,以隔离日志分离误差对整体WER性能的影响。