In this paper, we propose a speaker verification method by an Attentive Multi-scale Convolutional Recurrent Network (AMCRN). The proposed AMCRN can acquire both local spatial information and global sequential information from the input speech recordings. In the proposed method, logarithm Mel spectrum is extracted from each speech recording and then fed to the proposed AMCRN for learning speaker embedding. Afterwards, the learned speaker embedding is fed to the back-end classifier (such as cosine similarity metric) for scoring in the testing stage. The proposed method is compared with state-of-the-art methods for speaker verification. Experimental data are three public datasets that are selected from two large-scale speech corpora (VoxCeleb1 and VoxCeleb2). Experimental results show that our method exceeds baseline methods in terms of equal error rate and minimal detection cost function, and has advantages over most of baseline methods in terms of computational complexity and memory requirement. In addition, our method generalizes well across truncated speech segments with different durations, and the speaker embedding learned by the proposed AMCRN has stronger generalization ability across two back-end classifiers.
翻译:本文提出了一种基于注意力机制的多尺度卷积循环网络(AMCRN)的说话人验证方法。所提出的AMCRN能够从输入语音记录中同时获取局部空间信息和全局序列信息。在该方法中,首先从每条语音记录中提取对数梅尔频谱,然后将其输入到所提出的AMCRN中以学习说话人嵌入。随后,在测试阶段,将学习到的说话人嵌入输入到后端分类器(例如余弦相似度度量)中进行评分。我们将所提出的方法与当前最先进的说话人验证方法进行了比较。实验数据选自两个大规模语音语料库(VoxCeleb1和VoxCeleb2)中的三个公开数据集。实验结果表明,我们的方法在等错误率和最小检测代价函数方面优于基线方法,并且在计算复杂度和内存需求方面优于大多数基线方法。此外,我们的方法对不同时长的截断语音片段具有良好的泛化能力,并且通过所提出的AMCRN学习到的说话人嵌入在两个后端分类器上具有更强的泛化能力。