Acoustic echo cancellation (AEC) aims to remove interference signals while leaving near-end speech least distorted. As the indistinguishable patterns between near-end speech and interference signals, near-end speech can't be separated completely, causing speech distortion and interference signals residual. We observe that besides target positive information, e.g., ground-truth speech and features, the target negative information, such as interference signals and features, helps make pattern of target speech and interference signals more discriminative. Therefore, we present a novel AEC model encoder-decoder architecture with the guidance of negative information termed as CMNet. A collaboration module (CM) is designed to establish the correlation between the target positive and negative information in a learnable manner via three blocks: target positive, target negative, and interactive block. Experimental results demonstrate our CMNet achieves superior performance than recent methods.
翻译:声学回声消除(AEC)旨在去除干扰信号的同时最小化近端语音失真。由于近端语音与干扰信号存在不可区分的模式特征,近端语音无法被完全分离,导致语音失真和干扰信号残留。我们观察到,除目标正信息(如真实语音及其特征)外,目标负信息(如干扰信号及其特征)有助于使目标语音与干扰信号的模式更具区分性。为此,我们提出一种新颖的AEC模型——基于负信息引导的编解码器架构CMNet。通过目标正模块、目标负模块和交互模块三个组件,设计了一种可学习的协作模块(CM)来建立目标正负信息之间的关联。实验结果表明,我们的CMNet在性能上优于近期提出的方法。