In this work, we introduce metric learning (ML) to enhance the deep embedding learning for text-independent speaker verification (SV). Specifically, the deep speaker embedding network is trained with conventional cross entropy loss and auxiliary pair-based ML loss function. For the auxiliary ML task, training samples of a mini-batch are first arranged into pairs, then positive and negative pairs are selected and weighted through their own and relative similarities, and finally the auxiliary ML loss is calculated by the similarity of the selected pairs. To evaluate the proposed method, we conduct experiments on the Speaker in the Wild (SITW) dataset. The results demonstrate the effectiveness of the proposed method.
翻译:本文引入度量学习(Metric Learning, ML)以增强文本无关说话人验证(Text-independent Speaker Verification, SV)中的深度嵌入学习。具体而言,深度说话人嵌入网络采用传统交叉熵损失与辅助的基于成对样本的度量学习损失函数进行联合训练。对于辅助度量学习任务,首先将小批量训练样本配对,随后依据样本自身相似度及相对相似度选取并加权正负样本对,最终通过所选样本对的相似度计算辅助度量学习损失。为评估所提方法,我们在Speaker in the Wild (SITW)数据集上开展实验。结果表明所提方法具有有效性。