Diffusion models are a new class of generative models that have recently been applied to speech enhancement successfully. Previous works have demonstrated their superior performance in mismatched conditions compared to state-of-the art discriminative models. However, this was investigated with a single database for training and another one for testing, which makes the results highly dependent on the particular databases. Moreover, recent developments from the image generation literature remain largely unexplored for speech enhancement. These include several design aspects of diffusion models, such as the noise schedule or the reverse sampler. In this work, we systematically assess the generalization performance of a diffusion-based speech enhancement model by using multiple speech, noise and binaural room impulse response (BRIR) databases to simulate mismatched acoustic conditions. We also experiment with a noise schedule and a sampler that have not been applied to speech enhancement before. We show that the proposed system substantially benefits from using multiple databases for training, and achieves superior performance compared to state-of-the-art discriminative models in both matched and mismatched conditions. We also show that a Heun-based sampler achieves superior performance at a smaller computational cost compared to a sampler commonly used for speech enhancement.
翻译:扩散模型是一类新兴的生成模型,近年来已被成功应用于语音增强领域。已有研究表明,与最先进的判别模型相比,扩散模型在不匹配条件下展现出更优越的性能。然而,这些研究均使用单一数据库进行训练、另一数据库进行测试,导致结果高度依赖于特定数据库。此外,图像生成领域的最新进展(包括噪声调度策略和反向采样器等扩散模型设计要素)在语音增强中尚未得到充分探索。本研究通过使用多个语音、噪声和双耳房间冲激响应(BRIR)数据库模拟不匹配声学条件,系统评估了基于扩散模型的语音增强方法的泛化性能。我们还实验了两种此前未应用于语音增强的噪声调度策略和采样器。实验表明,所提系统因采用多数据库训练而显著受益,在匹配与不匹配条件下均优于最先进的判别模型。同时,与语音增强领域常用的采样器相比,基于Heun的采样器能以更小的计算成本实现更优性能。