We present in this paper an informed single-channel dereverberation method based on conditional generation with diffusion models. With knowledge of the room impulse response, the anechoic utterance is generated via reverse diffusion using a measurement consistency criterion coupled with a neural network that represents the clean speech prior. The proposed approach is largely more robust to measurement noise compared to a state-of-the-art informed single-channel dereverberation method, especially for non-stationary noise. Furthermore, we compare to other blind dereverberation methods using diffusion models and show superiority of the proposed approach for large reverberation times. We motivate the presented algorithm by introducing an extension for blind dereverberation allowing joint estimation of the room impulse response and anechoic speech. Audio samples and code can be found online (https://uhh.de/inf-sp-derev-dps).
翻译:本文提出一种基于扩散模型条件生成的信息型单通道去混响方法。在已知房间脉冲响应的条件下,通过结合用于表征纯净语音先验的神经网络与测量一致性准则,利用反向扩散过程生成消声语音。相较于现有的信息型单通道去混响方法,本方法对测量噪声具有更强的鲁棒性,尤其适用于非平稳噪声场景。此外,与基于扩散模型的其他盲去混响方法相比,本方法在长混响时间条件下展现出更优性能。我们通过引入一种可联合估计房间脉冲响应与消声语音的盲去混响扩展算法,对本文提出的方法进行了理论延伸。相关音频样本与代码已公开于在线资源库(https://uhh.de/inf-sp-derev-dps)。