For the task of speech separation, previous study usually treats multi-channel and single-channel scenarios as two research tracks with specialized solutions developed respectively. Instead, we propose a simple and unified architecture - DasFormer (Deep alternating spectrogram transFormer) to handle both of them in the challenging reverberant environments. Unlike frame-wise sequence modeling, each TF-bin in the spectrogram is assigned with an embedding encoding spectral and spatial information. With such input, DasFormer is then formed by multiple repetition of simple blocks each of which integrates 1) two multi-head self-attention (MHSA) modules alternately processing within each frequency bin & temporal frame of the spectrogram 2) MBConv before each MHSA for modeling local features on the spectrogram. Experiments show that DasFormer has a powerful ability to model the time-frequency representation, whose performance far exceeds the current SOTA models in multi-channel speech separation, and also achieves single-channel SOTA in the more challenging yet realistic reverberation scenario.
翻译:针对语音分离任务,以往研究通常将多通道和单通道场景视为两条独立的研究路线,并分别开发专用解决方案。取而代之,我们提出一种简洁统一的架构——DasFormer(深度交替频谱变换器),以在具有挑战性的混响环境中同时处理这两种场景。与基于帧的序列建模不同,频谱图中的每个时频单元被赋予一个同时编码频谱和空间信息的嵌入表示。基于这种输入,DasFormer由多个重复的简单模块堆叠而成,每个模块整合:1)两个多头自注意力模块,交替处理频谱图内每个频率仓和时域帧;2)在每个多头自注意力模块前插入MBConv,用于建模频谱图的局部特征。实验表明,DasFormer具备强大的时频表征建模能力,其在多通道语音分离任务中的性能大幅超越当前最优模型,并在更具挑战性与真实性的混响场景中同样实现了单通道分离的最优结果。