Sound source localisation is used in many consumer electronics devices, to help isolate audio from individual speakers and to reject noise. Localization is frequently accomplished by "beamforming" algorithms, which combine microphone audio streams to improve received signal power from particular incident source directions. Beamforming algorithms generally use knowledge of the frequency components of the audio source, along with the known microphone array geometry, to analytically phase-shift microphone streams before combining them. A dense set of band-pass filters is often used to obtain known-frequency "narrowband" components from wide-band audio streams. These approaches achieve high accuracy, but state of the art narrowband beamforming algorithms are computationally demanding, and are therefore difficult to integrate into low-power IoT devices. We demonstrate a novel method for sound source localisation in arbitrary microphone arrays, designed for efficient implementation in ultra-low-power spiking neural networks (SNNs). We use a novel short-time Hilbert transform (STHT) to remove the need for demanding band-pass filtering of audio, and introduce a new accompanying method for audio encoding with spiking events. Our beamforming and localisation approach achieves state-of-the-art accuracy for SNN methods, and comparable with traditional non-SNN super-resolution approaches. We deploy our method to low-power SNN audio inference hardware, and achieve much lower power consumption compared with super-resolution methods. We demonstrate that signal processing approaches can be co-designed with spiking neural network implementations to achieve high levels of power efficiency. Our new Hilbert-transform-based method for beamforming promises to also improve the efficiency of traditional DSP-based signal processing.
翻译:音频源定位技术广泛应用于消费电子设备中,用于分离单个说话人的音频信号并抑制噪声。传统上,定位通过"波束成形"算法实现,该算法组合多个麦克风音频流,以增强特定入射方向信号接收功率。波束成形算法通常利用音频源的频率分量信息及已知的麦克风阵列几何结构,在组合前对麦克风流进行解析相移。为从宽带音频流中获取已知频率的"窄带"分量,常需使用密集带通滤波器组。这些方法虽能实现高精度,但当前最先进的窄带波束成形算法计算需求极高,难以集成至低功耗物联网设备中。本文提出一种适用于任意麦克风阵列的新型音频源定位方法,专为超低功耗脉冲神经网络(SNN)的高效实现而设计。我们采用新型短时希尔伯特变换(STHT)替代计算密集的带通滤波音频处理,并引入配套的脉冲事件音频编码方法。所提出的波束成形与定位方法在SNN方法中达到当前最优精度,可与传统非SNN超分辨率方法相媲美。将该方法部署至低功耗SNN音频推理硬件后,功耗较超分辨率方法显著降低。实验证明,信号处理方法与脉冲神经网络实现协同设计可达成高能效水平。这项基于希尔伯特变换的新型波束成形方法,也有望提升传统DSP信号处理的效率。