This paper focuses on leveraging deep representation learning (DRL) for speech enhancement (SE). In general, the performance of the deep neural network (DNN) is heavily dependent on the learning of data representation. However, the DRL's importance is often ignored in many DNN-based SE algorithms. To obtain a higher quality enhanced speech, we propose a two-stage DRL-based SE method through adversarial training. In the first stage, we disentangle different latent variables because disentangled representations can help DNN generate a better enhanced speech. Specifically, we use the $\beta$-variational autoencoder (VAE) algorithm to obtain the speech and noise posterior estimations and related representations from the observed signal. However, since the posteriors and representations are intractable and we can only apply a conditional assumption to estimate them, it is difficult to ensure that these estimations are always pretty accurate, which may potentially degrade the final accuracy of the signal estimation. To further improve the quality of enhanced speech, in the second stage, we introduce adversarial training to reduce the effect of the inaccurate posterior towards signal reconstruction and improve the signal estimation accuracy, making our algorithm more robust for the potentially inaccurate posterior estimations. As a result, better SE performance can be achieved. The experimental results indicate that the proposed strategy can help similar DNN-based SE algorithms achieve higher short-time objective intelligibility (STOI), perceptual evaluation of speech quality (PESQ), and scale-invariant signal-to-distortion ratio (SI-SDR) scores. Moreover, the proposed algorithm can also outperform recent competitive SE algorithms.
翻译:本文聚焦于利用深度表示学习(DRL)进行语音增强(SE)。一般而言,深度神经网络(DNN)的性能高度依赖于数据表示的学习。然而,在许多基于DNN的SE算法中,DRL的重要性常被忽视。为获得更高质量的增强语音,我们通过对抗训练提出了一种两阶段DRL-SE方法。在第一阶段,我们解耦不同的潜在变量,因为解耦表示有助于DNN生成更优的增强语音。具体而言,我们采用$\beta$-变分自编码器(VAE)算法,从观测信号中获取语音和噪声的后验估计及相关表示。然而,由于后验分布和表示难以直接计算,我们只能通过条件假设进行估计,因此难以保证这些估计始终精确,这可能潜在降低信号估计的最终准确性。为进一步提升增强语音质量,在第二阶段,我们引入对抗训练,以降低不精确后验对信号重建的影响,并提高信号估计精度,使算法对潜在不精确的后验估计更具鲁棒性。最终,可获得更优的SE性能。实验结果表明,所提策略能帮助类似的基于DNN的SE算法获得更高的短时客观可懂度(STOI)、感知语音质量评估(PESQ)和尺度不变信噪比(SI-SDR)分数。此外,所提算法性能也优于近期具有竞争力的SE算法。