Speech deepfake source verification systems aims to determine whether two synthetic speech utterances originate from the same source generator, often assuming that the resulting source embeddings are independent of speaker traits. However, this assumption remains unverified. In this paper, we first investigate the impact of speaker factors on source verification. We propose a speaker-disentangled metric learning (SDML) framework incorporating two novel loss functions. The first leverages Chebyshev polynomial to mitigate gradient instability during disentanglement optimization. The second projects source and speaker embeddings into hyperbolic space, leveraging Riemannian metric distances to reduce speaker information and learn more discriminative source features. Experimental results on MLAAD benchmark, evaluated under four newly proposed protocols designed for source-speaker disentanglement scenarios, demonstrate the effectiveness of SDML framework. The code, evaluation protocols and demo website are available at https://github.com/xxuan-acoustics/RiemannSD-Net.
翻译:语音深度伪造声源验证系统旨在判断两段合成语音是否来自同一生成源,通常假设提取的源嵌入特征与说话者特征无关。然而,这一假设尚未得到验证。本文首先研究了说话者因素对声源验证的影响,提出了一种说话者解耦度量学习(SDML)框架,该框架包含两个新型损失函数:第一个利用切比雪夫多项式缓解解耦优化过程中的梯度不稳定性;第二个将源嵌入和说话者嵌入投影至双曲空间,利用黎曼度量距离来减少说话者信息,并学习更具判别性的源特征。在MLAAD基准数据集上,针对源-说话者解耦场景设计的四种新评估协议实验结果验证了SDML框架的有效性。代码、评估协议及演示网站详见https://github.com/xxuan-acoustics/RiemannSD-Net。