Self-supervised learned models have been found to be very effective for certain speech tasks such as automatic speech recognition, speaker identification, keyword spotting and others. While the features are undeniably useful in speech recognition and associated tasks, their utility in speech enhancement systems is yet to be firmly established, and perhaps not properly understood. In this paper, we investigate the uses of SSL representations for single-channel speech enhancement in challenging conditions and find that they add very little value for the enhancement task. Our constraints are designed around on-device real-time speech enhancement -- model is causal, the compute footprint is small. Additionally, we focus on low SNR conditions where such models struggle to provide good enhancement. In order to systematically examine how SSL representations impact performance of such enhancement models, we propose a variety of techniques to utilize these embeddings which include different forms of knowledge-distillation and pre-training.
翻译:自监督学习模型已被证实对自动语音识别、说话人识别、关键词唤醒等语音任务具有显著效果。尽管这些特征在语音识别及相关任务中的价值毋庸置疑,但其在语音增强系统中的实用性尚未得到充分验证,甚至可能未被正确理解。本文探究了自监督学习表征在复杂条件下单通道语音增强中的应用,发现其对增强任务的增益极为有限。我们基于设备端实时语音增强的约束条件进行设计——模型需具备因果性,且计算开销较小。此外,我们重点关注低信噪比场景,此类模型在此条件下难以实现优质增强。为系统分析自监督学习表征对增强模型性能的影响机制,我们提出了多种嵌入表示利用策略,涵盖不同形式的知识蒸馏与预训练方法。