Knowledge distillation (KD) is used to enhance automatic speaker verification performance by ensuring consistency between large teacher networks and lightweight student networks at the embedding level or label level. However, the conventional label-level KD overlooks the significant knowledge from non-target speakers, particularly their classification probabilities, which can be crucial for automatic speaker verification. In this paper, we first demonstrate that leveraging a larger number of training non-target speakers improves the performance of automatic speaker verification models. Inspired by this finding about the importance of non-target speakers' knowledge, we modified the conventional label-level KD by disentangling and emphasizing the classification probabilities of non-target speakers during knowledge distillation. The proposed method is applied to three different student model architectures and achieves an average of 13.67% improvement in EER on the VoxCeleb dataset compared to embedding-level and conventional label-level KD methods.
翻译:知识蒸馏(KD)通过确保大教师网络与轻量级学生网络在嵌入层或标签层的一致性,被用于提升自动说话人验证性能。然而,传统标签层KD忽略了非目标说话人的关键知识,特别是其分类概率,这些知识对自动说话人验证至关重要。本文首先证明利用更多训练非目标说话人可提升自动说话人验证模型的性能。受此关于非目标说话人知识重要性的发现启发,我们改进传统标签层KD,在知识蒸馏过程中分离并强调非目标说话人的分类概率。所提方法应用于三种不同学生模型架构,在VoxCeleb数据集上相较于嵌入层和传统标签层KD方法实现了平均13.67%的等错误率(EER)改进。