The accuracy of automated speaker recognition is negatively impacted by change in emotions in a person's speech. In this paper, we hypothesize that speaker identity is composed of various vocal style factors that may be learned from unlabeled data and re-combined using a neural network to generate a holistic speaker identity representation for affective scenarios. In this regard, we propose the E-Vector architecture, composed of a 1-D CNN for learning speaker identity features and a vocal style factorization technique for determining vocal styles. Experiments conducted on the MSP-Podcast dataset demonstrate that the proposed architecture improves state-of-the-art speaker recognition accuracy in the affective domain over baseline ECAPA-TDNN speaker recognition models. For instance, the true match rate at a false match rate of 1% improves from 27.6% to 46.2%.
翻译:自动说话人识别的准确性会受到语音中情绪变化的负面影响。本文假设说话人身份由多种可从无标签数据中学习的发声风格因子组成,并且可以通过神经网络重新组合,生成情感场景下的整体说话人身份表征。为此,我们提出E-Vector架构,该架构包含一个用于学习说话人身份特征的一维卷积神经网络,以及一种用于确定发声风格的分解技术。在MSP-Podcast数据集上进行的实验表明,所提出的架构在情感域中的说话人识别准确率优于基线ECAPA-TDNN模型。例如,在1%的错误接受率下,正确匹配率从27.6%提升至46.2%。