This paper describes our submission to the Second Clarity Enhancement Challenge (CEC2), which consists of target speech enhancement for hearing-aid (HA) devices in noisy-reverberant environments with multiple interferers such as music and competing speakers. Our approach builds upon the powerful iterative neural/beamforming enhancement (iNeuBe) framework introduced in our recent work, and this paper extends it for target speaker extraction. We therefore name the proposed approach as iNeuBe-X, where the X stands for extraction. To address the challenges encountered in the CEC2 setting, we introduce four major novelties: (1) we extend the state-of-the-art TF-GridNet model, originally designed for monaural speaker separation, for multi-channel, causal speech enhancement, and large improvements are observed by replacing the TCNDenseNet used in iNeuBe with this new architecture; (2) we leverage a recent dual window size approach with future-frame prediction to ensure that iNueBe-X satisfies the 5 ms constraint on algorithmic latency required by CEC2; (3) we introduce a novel speaker-conditioning branch for TF-GridNet to achieve target speaker extraction; (4) we propose a fine-tuning step, where we compute an additional loss with respect to the target speaker signal compensated with the listener audiogram. Without using external data, on the official development set our best model reaches a hearing-aid speech perception index (HASPI) score of 0.942 and a scale-invariant signal-to-distortion ratio improvement (SI-SDRi) of 18.8 dB. These results are promising given the fact that the CEC2 data is extremely challenging (e.g., on the development set the mixture SI-SDR is -12.3 dB). A demo of our submitted system is available at WAVLab CEC2 demo.
翻译:本文介绍了我们在第二届听力增强挑战赛(CEC2)中的提交方案,该方案针对助听器(HA)设备在噪声混响环境中,面对音乐和竞争说话者等多干扰源时的目标语音增强问题。我们的方法基于近期工作中提出的强大的迭代神经/波束成形增强(iNeuBe)框架,并在本文中将其扩展至目标说话人提取。因此,我们将所提方法命名为 iNeuBe-X,其中 X 代表提取。为应对 CEC2 设定中的挑战,我们引入了四项主要创新:(1)将最初为单声道说话人分离设计的最先进 TF-GridNet 模型扩展至多声道因果语音增强,通过用此新架构替换 iNeuBe 中使用的 TCNDenseNet,我们观察到了显著的性能提升;(2)采用最近的具有未来帧预测的双窗口尺寸方法,确保 iNeuBe-X 满足 CEC2 要求的 5 毫秒算法延迟限制;(3)为 TF-GridNet 引入了一种新颖的说话人条件分支,以实现目标说话人提取;(4)提出了一种微调步骤,其中我们根据经听者听力图补偿后的目标说话人信号计算额外损失。在不使用外部数据的情况下,在官方开发集上,我们的最佳模型达到了助听器语音感知指数(HASPI)0.942 和尺度不变信噪比改善(SI-SDRi)18.8 dB。考虑到 CEC2 数据极具挑战性(例如,开发集中混合信号的 SI-SDR 为 -12.3 dB),这些结果令人鼓舞。我们提交系统的演示可在 WAVLab CEC2 演示页面获取。