Training robust speaker verification systems without speaker labels has long been a challenging task. Previous studies observed a large performance gap between self-supervised and fully supervised methods. In this paper, we apply a non-contrastive self-supervised learning framework called DIstillation with NO labels (DINO) and propose two regularization terms applied to embeddings in DINO. One regularization term guarantees the diversity of the embeddings, while the other regularization term decorrelates the variables of each embedding. The effectiveness of various data augmentation techniques are explored, on both time and frequency domain. A range of experiments conducted on the VoxCeleb datasets demonstrate the superiority of the regularized DINO framework in speaker verification. Our method achieves the state-of-the-art speaker verification performance under a single-stage self-supervised setting on VoxCeleb.
翻译:长期以来,在无说话人标签条件下训练鲁棒的说话人验证系统一直是一项具有挑战性的任务。先前的研究观察到自监督方法与全监督方法之间存在显著的性能差距。本文采用了一种名为无标签蒸馏(DINO)的非对比自监督学习框架,并针对该框架中的嵌入表示提出了两项正则化约束。其中一项正则化项保证了嵌入表示的多样性,另一项则通过解耦各嵌入内部的变量相关性来优化。同时,本文在时域与频域上探索了多种数据增强技术的有效性。在VoxCeleb数据集上进行的一系列实验表明,经过正则化改进的DINO框架在说话人验证任务中具有显著优势。所提方法在VoxCeleb单阶段自监督设置下达到了当前最优的说话人验证性能。