Supervised learning based on a deep neural network recently has achieved substantial improvement on speech enhancement. Denoising networks learn mapping from noisy speech to clean one directly, or to a spectrum mask which is the ratio between clean and noisy spectra. In either case, the network is optimized by minimizing mean square error (MSE) between ground-truth labels and time-domain or spectrum output. However, existing schemes have either of two critical issues: spectrum and metric mismatches. The spectrum mismatch is a well known issue that any spectrum modification after short-time Fourier transform (STFT), in general, cannot be fully recovered after inverse short-time Fourier transform (ISTFT). The metric mismatch is that a conventional MSE metric is sub-optimal to maximize our target metrics, signal-to-distortion ratio (SDR) and perceptual evaluation of speech quality (PESQ). This paper presents a new end-to-end denoising framework with the goal of joint SDR and PESQ optimization. First, the network optimization is performed on the time-domain signals after ISTFT to avoid spectrum mismatch. Second, two loss functions which have improved correlations with SDR and PESQ metrics are proposed to minimize metric mismatch. The experimental result showed that the proposed denoising scheme significantly improved both SDR and PESQ performance over the existing methods.
翻译:基于深度神经网络的监督学习近年来在语音增强领域取得了显著进展。去噪网络可直接学习从带噪语音到干净语音的映射,或学习表征干净谱与带噪谱比值的频谱掩码。无论采用何种方式,网络均通过最小化真实标签与时域或频域输出之间的均方误差(MSE)进行优化。然而,现有方案存在频谱失配与度量失配两大关键问题。频谱失配是指:经短时傅里叶变换(STFT)后的频谱修改通常无法通过逆短时傅里叶变换(ISTFT)完全恢复。度量失配则体现为:传统MSE指标在最大化目标度量——信号失真比(SDR)与语音感知评估(PESQ)时并非最优。本文提出一种新的端到端去噪框架,旨在联合优化SDR与PESQ。首先,网络优化在ISTFT后的时域信号上执行,以避免频谱失配;其次,提出两种与SDR和PESQ指标具有更强相关性的损失函数,以最小化度量失配。实验结果表明,所提去噪方案在SDR和PESQ性能上均较现有方法有显著提升。