Beamforming for multichannel speech enhancement relies on the estimation of spatial characteristics of the acoustic scene. In its simplest form, the delay-and-sum beamformer (DSB) introduces a time delay to all channels to align the desired signal components for constructive superposition. Recent investigations of neural spatiospectral filtering revealed that these filters can be characterized by a beampattern similar to one of traditional beamformers, which shows that artificial neural networks can learn and explicitly represent spatial structure. Using the Complex-valued Spatial Autoencoder (COSPA) as an exemplary neural spatiospectral filter for multichannel speech enhancement, we investigate where and how such networks represent spatial information. We show via clustering that for COSPA the spatial information is represented by the features generated by a gated recurrent unit (GRU) layer that has access to all channels simultaneously and that these features are not source -- but only direction of arrival-dependent.
翻译:多通道语音增强的波束形成依赖于对声学场景空间特性的估计。其最基本形式——延迟求和波束形成器通过对所有通道施加时间延迟,使期望信号分量对齐以进行相长叠加。近期对神经空间谱滤波的研究揭示,此类滤波器呈现与经典波束形成器相似的波束方向图,表明人工神经网络能够学习并显式表达空间结构。本文以复值空间自编码器(COSPA)作为多通道语音增强的神经空间谱滤波示例,探究此类网络表示空间信息的位置与方式。通过聚类分析表明,在COSPA中,空间信息由能够同时访问所有通道的门控循环单元层生成的特征所表征,且这些特征不依赖于声源身份,仅与到达方向有关。