Noise suppression models running in production environments are commonly trained on publicly available datasets. However, this approach leads to regressions due to the lack of training/testing on representative customer data. Moreover, due to privacy reasons, developers cannot listen to customer content. This `ears-off' situation motivates augmenting existing datasets in a privacy-preserving manner. In this paper, we present Aura, a solution to make existing noise suppression test sets more challenging and diverse while being sample efficient. Aura is `ears-off' because it relies on a feature extractor and a metric of speech quality, DNSMOS P.835, both pre-trained on data obtained from public sources. As an application of Aura, we augment the INTERSPEECH 2021 DNS challenge by sampling audio files from a new batch of data of 20K clean speech clips from Librivox mixed with noise clips obtained from AudioSet. Aura makes the existing benchmark test set harder by 0.27 in DNSMOS P.835 OVLR (7%), 0.64 harder in DNSMOS P.835 SIG (16%), increases diversity by 31%, and achieves a 26% improvement in Spearman's rank correlation coefficient (SRCC) compared to random sampling. Finally, we open-source Aura to stimulate research of test set development.
翻译:在生产环境中运行的噪声抑制模型通常基于公开数据集进行训练。然而,由于缺乏代表性客户数据的训练与测试,这会导致性能退化。此外,出于隐私原因,开发者无法听取客户内容。这种"无耳"状态促使我们以保护隐私的方式增强现有数据集。本文提出Aura方案,该方案能在保持样本效率的前提下,使现有噪声抑制测试集更具挑战性和多样性。Aura实现"无耳"机制的关键在于依赖特征提取器与语音质量指标DNSMOS P.835——两者均基于公开来源数据预训练。作为Aura的应用案例,我们从Librivox的2万条清洁语音片段与AudioSet噪声片段混合的新数据批次中采样音频文件,增强INTERSPEECH 2021 DNS挑战数据集。相比随机采样,Aura使现有基准测试集在DNSMOS P.835 OVLR指标上难度提升0.27(7%),DNSMOS P.835 SIG指标提升0.64(16%),多样性提高31%,斯皮尔曼秩相关系数(SRCC)提升26%。最后,我们开源Aura以促进测试集开发相关研究。