Supervised learning based on a deep neural network recently has achieved substantial improvement on speech enhancement. Denoising networks learn mapping from noisy speech to clean one directly, or to a spectrum mask which is the ratio between clean and noisy spectra. In either case, the network is optimized by minimizing mean square error (MSE) between ground-truth labels and time-domain or spectrum output. However, existing schemes have either of two critical issues: spectrum and metric mismatches. The spectrum mismatch is a well known issue that any spectrum modification after short-time Fourier transform (STFT), in general, cannot be fully recovered after inverse short-time Fourier transform (ISTFT). The metric mismatch is that a conventional MSE metric is sub-optimal to maximize our target metrics, signal-to-distortion ratio (SDR) and perceptual evaluation of speech quality (PESQ). This paper presents a new end-to-end denoising framework with the goal of joint SDR and PESQ optimization. First, the network optimization is performed on the time-domain signals after ISTFT to avoid spectrum mismatch. Second, two loss functions which have improved correlations with SDR and PESQ metrics are proposed to minimize metric mismatch. The experimental result showed that the proposed denoising scheme significantly improved both SDR and PESQ performance over the existing methods.
翻译:近年来,基于深度神经网络的监督学习在语音增强领域取得了显著进展。降噪网络可直接学习从带噪语音到纯净语音的映射,或学习表征纯净与带噪频谱之比的频谱掩模。无论采用何种方式,网络均通过最小化真实标签与时域或频谱输出之间的均方误差(MSE)进行优化。然而,现有方案存在频谱失配与度量失配两大关键问题。频谱失配是指:经短时傅里叶变换(STFT)后,任何频谱修正通常无法通过逆短时傅里叶变换(ISTFT)完全恢复;度量失配则指传统MSE度量并非最大化目标度量——信号失真比(SDR)与语音质量感知评价(PESQ)的最优选择。本文提出一种全新的端到端降噪框架,旨在联合优化SDR与PESQ指标。首先,网络优化在ISTFT后的时域信号上执行,以避免频谱失配问题;其次,提出两种与SDR和PESQ度量具有更高相关性的损失函数,以最小化度量失配。实验结果表明,所提降噪方案较现有方法显著提升了SDR与PESQ性能。