Self-supervised speech representation learning (S3RL) is revolutionizing the way we leverage the ever-growing availability of data. While S3RL related studies typically use large models, we employ light-weight networks to comply with tight memory of compute-constrained devices. We demonstrate the effectiveness of S3RL on a keyword-spotting (KS) problem by using transformers with 330k parameters and propose a mechanism to enhance utterance-wise distinction, which proves crucial for improving performance on classification tasks. On the Google speech commands v2 dataset, the proposed method applied to the Auto-Regressive Predictive Coding S3RL led to a 1.2% accuracy improvement compared to training from scratch. On an in-house KS dataset with four different keywords, it provided 6% to 23.7% relative false accept improvement at fixed false reject rate. We argue this demonstrates the applicability of S3RL approaches to light-weight models for KS and confirms S3RL is a powerful alternative to traditional supervised learning for resource-constrained applications.
翻译:自监督语音表示学习(S3RL)正彻底改变我们利用日益增长的数据可用性的方式。尽管S3RL相关研究通常采用大型模型,但我们使用轻量级网络以满足计算受限设备的紧凑内存需求。我们通过采用33万参数的Transformer在关键词检测(KS)问题中展示了S3RL的有效性,并提出了一种增强话语间区分度的机制,这对提升分类任务性能至关重要。在Google语音指令v2数据集上,将所提方法应用于自回归预测编码S3RL后,相比从头训练实现了1.2%的准确率提升。在包含四个不同关键词的内部KS数据集上,在固定误拒率条件下,相对误接受率改善了6%至23.7%。我们论证了S3RL方法在轻量级KS模型中的适用性,并确认S3RL是资源受限应用中传统监督学习的强大替代方案。