Speech utterances recorded under differing conditions exhibit varying degrees of confidence in their embedding estimates, i.e., uncertainty, even if they are extracted using the same neural network. This paper aims to incorporate the uncertainty estimate produced in the xi-vector network front-end with a probabilistic linear discriminant analysis (PLDA) back-end scoring for speaker verification. To achieve this we derive a posterior covariance matrix, which measures the uncertainty, from the frame-wise precisions to the embedding space. We propose a log-likelihood ratio function for the PLDA scoring with the uncertainty propagation. We also propose to replace the length normalization pre-processing technique with a length scaling technique for the application of uncertainty propagation in the back-end. Experimental results on the VoxCeleb-1, SITW test sets as well as a domain-mismatched CNCeleb1-E set show the effectiveness of the proposed techniques with 14.5%-41.3% EER reductions and 4.6%-25.3% minDCF reductions.
翻译:在不同条件下录制的语音话语,即使使用相同的神经网络进行嵌入提取,其嵌入估计的置信度(即不确定性)也会表现出不同程度的差异。本文旨在将xi向量网络前端生成的不确定性估计与概率线性判别分析(PLDA)后端的评分相结合,用于说话人验证。为实现这一目标,我们从帧级精度推导出后验协方差矩阵,该矩阵用于衡量嵌入空间中的不确定性。我们提出了一个带有不确定性传播的PLDA评分对数似然比函数。同时,我们建议用长度缩放技术替代长度归一化预处理技术,以在后端实现不确定性传播的应用。在VoxCeleb-1、SITW测试集以及域不匹配的CNCeleb1-E集上的实验结果表明,所提出的技术有效实现了14.5%-41.3%的等错误率(EER)降低和4.6%-25.3%的最小检测成本函数(minDCF)降低。