Automated speaker identification (SID) is a crucial step for the personalization of a wide range of speech-enabled services. Typical SID systems use a symmetric enrollment-verification framework with a single model to derive embeddings both offline for voice profiles extracted from enrollment utterances, and online from runtime utterances. Due to the distinct circumstances of enrollment and runtime, such as different computation and latency constraints, several applications would benefit from an asymmetric enrollment-verification framework that uses different models for enrollment and runtime embedding generation. To support this asymmetric SID where each of the two models can be updated independently, we propose using a lightweight neural network to map the embeddings from the two independent models to a shared speaker embedding space. Our results show that this approach significantly outperforms cosine scoring in a shared speaker logit space for models that were trained with a contrastive loss on large datasets with many speaker identities. This proposed Neural Embedding Speaker Space Alignment (NESSA) combined with an asymmetric update of only one of the models delivers at least 60% of the performance gain achieved by updating both models in the standard symmetric SID approach.
翻译:自动说话人识别(SID)是众多语音服务个性化中的关键步骤。典型的SID系统采用对称的注册-验证框架,使用单一模型同时从注册语句中离线提取语音特征嵌入,以及从运行时语句中在线提取嵌入。由于注册和运行时的条件不同,例如计算和延迟约束的差异,许多应用将受益于非对称的注册-验证框架,该框架使用不同模型分别生成注册和运行时的嵌入。为支持这种两个模型可独立更新的非对称SID,我们提出使用轻量级神经网络将来自两个独立模型的嵌入映射到共享的说话人嵌入空间。结果表明,对于在大规模多说话人数据集上使用对比损失训练的模型,该方法在共享说话人logit空间中的表现显著优于余弦评分。我们提出的神经嵌入说话人空间对齐(NESSA)结合仅更新其中一个模型的非对称方法,至少能达到标准对称SID方法中同时更新两个模型所获得性能增益的60%。