ECAPA-TDNN is currently the most popular TDNN-series model for speaker verification, which refreshed the state-of-the-art(SOTA) performance of TDNN models. However, one-dimensional convolution has a global receptive field over the feature channel. It destroys the time-frequency relevance of the spectrogram. Besides, as ECAPA-TDNN only has five layers, a much shallower structure compared to ResNet restricts the capability to generate deep representations. To further improve ECAPA-TDNN, we propose a progressive channel fusion strategy that splits the spectrogram across the feature channel and gradually expands the receptive field through the network. Secondly, we enlarge the model by extending the depth and adding branches. Our proposed model achieves EER with 0.718 and minDCF(0.01) with 0.0858 on vox1o, relatively improved 16.1\% and 19.5\% compared with ECAPA-TDNN-large.
翻译:摘要:ECAPA-TDNN是当前说话人确认领域最流行的TDNN系列模型,它刷新了TDNN模型的最优性能。然而,一维卷积在特征通道上具有全局感受野,这会破坏语谱图的时频相关性。此外,ECAPA-TDNN仅包含五层网络,相较于ResNet更浅的结构限制了其生成深层表示的能力。为进一步改进ECAPA-TDNN,我们提出一种渐进通道融合策略:沿特征通道分割语谱图,并通过网络逐步扩展感受野。其次,我们通过加深网络深度并添加分支来扩大模型规模。在Vox1-O数据集上,本文提出的模型实现了0.718的等错误率和0.0858的minDCF(0.01),相比ECAPA-TDNN-large分别相对提升16.1%和19.5%。