Diffusion models are a new class of generative models that have recently been applied to speech enhancement successfully. Previous works have demonstrated their superior performance in mismatched conditions compared to state-of-the art discriminative models. However, this was investigated with a single database for training and another one for testing, which makes the results highly dependent on the particular databases. Moreover, recent developments from the image generation literature remain largely unexplored for speech enhancement. These include several design aspects of diffusion models, such as the noise schedule or the reverse sampler. In this work, we systematically assess the generalization performance of a diffusion-based speech enhancement model by using multiple speech, noise and binaural room impulse response (BRIR) databases to simulate mismatched acoustic conditions. We also experiment with a noise schedule and a sampler that have not been applied to speech enhancement before. We show that the proposed system substantially benefits from using multiple databases for training, and achieves superior performance compared to state-of-the-art discriminative models in both matched and mismatched conditions. We also show that a Heun-based sampler achieves superior performance at a smaller computational cost compared to a sampler commonly used for speech enhancement.
翻译:扩散模型是一类新型生成模型,近期已成功应用于语音增强领域。先前研究表明,与当前最先进的判别模型相比,扩散模型在失配条件下展现出更优性能。然而,这些研究仅使用单一数据库进行训练、另一数据库进行测试,导致结果高度依赖于特定数据库。此外,图像生成领域的最新进展(包括扩散模型的噪声调度的设计参数、反向采样器等)尚未在语音增强中得到充分探索。本研究通过使用多组语音、噪声及双耳房间脉冲响应(BRIR)数据库模拟失配声学条件,系统评估了基于扩散的语音增强模型的泛化性能。我们同时实验了此前未应用于语音增强的噪声调度方法与采样器。实验表明,所提系统显著受益于多数据库联合训练,并在匹配与失配条件下均取得了优于最先进判别模型的性能。我们还发现,基于Heun的采样器在降低计算成本的同时,相较语音增强中常用的采样器实现了更优的性能表现。