Noise suppression (NS) models have been widely applied to enhance speech quality. Recently, Deep Learning-Based NS, which we denote as Deep Noise Suppression (DNS), became the mainstream NS method due to its excelling performance over traditional ones. However, DNS models face 2 major challenges for supporting the real-world applications. First, high-performing DNS models are usually large in size, causing deployment difficulties. Second, DNS models require extensive training data, including noisy audios as inputs and clean audios as labels. It is often difficult to obtain clean labels for training DNS models. We propose the use of knowledge distillation (KD) to resolve both challenges. Our study serves 2 main purposes. To begin with, we are among the first to comprehensively investigate mainstream KD techniques on DNS models to resolve the two challenges. Furthermore, we propose a novel Attention-Based-Compression KD method that outperforms all investigated mainstream KD frameworks on DNS task.
翻译:噪声抑制(NS)模型已被广泛用于提升语音质量。近年来,基于深度学习的NS方法(我们称之为深度噪声抑制,DNS)因其优于传统方法的性能而成为主流NS方法。然而,DNS模型在实际应用中面临两大挑战:首先,高性能DNS模型通常体积较大,导致部署困难;其次,DNS模型需要大量训练数据,包括作为输入的含噪音频和作为标签的纯净音频,而纯净标签往往难以获得。我们提出使用知识蒸馏(KD)来解决这两个挑战。本研究有两个主要目的:首先,我们是首批全面探讨主流KD技术在DNS模型上解决这两个挑战的研究者之一;其次,我们提出了一种新颖的基于注意力压缩的KD方法,在DNS任务上性能优于所有被研究的主流KD框架。