Self-supervised learning has been used to leverage unlabelled data, improving accuracy and generalisation of speech systems through the training of representation models. While many recent works have sought to produce effective representations across a variety of acoustic domains, languages, modalities and even simultaneous speakers, these studies have all been limited to single-channel audio recordings. This paper presents Spatial HuBERT, a self-supervised speech representation model that learns both acoustic and spatial information pertaining to a single speaker in a potentially noisy environment by using multi-channel audio inputs. Spatial HuBERT learns representations that outperform state-of-the-art single-channel speech representations on a variety of spatial downstream tasks, particularly in reverberant and noisy environments. We also demonstrate the utility of the representations learned by Spatial HuBERT on a speech localisation downstream task. Along with this paper, we publicly release a new dataset of 100 000 simulated first-order ambisonics room impulse responses.
翻译:自监督学习已被用于利用未标注数据,通过训练表示模型提升语音系统的准确性和泛化能力。尽管近期许多研究致力于在多种声学领域、语言、模态乃至同时说话的多个说话人场景中生成有效表示,但这些研究均局限于单通道音频记录。本文提出空间HuBERT,一种自监督语音表示模型,它通过采用多通道音频输入,学习潜在噪声环境中单个说话人的声学与空间信息。空间HuBERT学习的表示在多种空间下游任务中(尤其在混响和噪声环境中)优于当前最先进的单通道语音表示。我们还证明了空间HuBERT学习的表示在语音定位下游任务中的实用性。伴随本文,我们公开发布了一个包含10万组仿一阶声全息房间脉冲响应模拟数据集。