The time-delay neural network (TDNN) is one of the state-of-the-art models for text-independent speaker verification. However, it is difficult for conventional TDNN to capture global context that has been proven critical for robust speaker representations and long-duration speaker verification in many recent works. Besides, the common solutions, e.g., self-attention, have quadratic complexity for input tokens, which makes them computationally unaffordable when applied to the feature maps with large sizes in TDNN. To address these issues, we propose the Global Filter for TDNN, which applies log-linear complexity FFT/IFFT and a set of differentiable frequency-domain filters to efficiently model the long-term dependencies in speech. Besides, a dynamic filtering strategy, and a sparse regularization method are specially designed to enhance the performance of the global filter and prevent it from overfitting. Furthermore, we construct a dual-stream TDNN (DS-TDNN), which splits the basic channels for complexity reduction and employs the global filter to increase recognition performance. Experiments on Voxceleb and SITW databases show that the DS-TDNN achieves approximate 10% improvement with a decline over 28% and 15% in complexity and parameters compared with the ECAPA-TDNN. Besides, it has the best trade-off between efficiency and effectiveness compared with other popular baseline systems when facing long-duration speech. Finally, visualizations and a detailed ablation study further reveal the advantages of the DS-TDNN.
翻译:时延神经网络(TDNN)是文本无关说话人确认任务中最先进的模型之一。然而,传统TDNN难以捕获全局上下文信息——近期诸多研究已证明全局上下文对鲁棒说话人表征及长时说话人确认至关重要。此外,常规解决方案(如自注意力机制)对输入令牌具有二次复杂度,当应用于TDNN中大尺寸特征图时,其计算成本难以承受。为解决这些问题,我们提出适用于TDNN的全局滤波器,该滤波器采用对数线性复杂度的快速傅里叶变换/逆快速傅里叶变换及一组可微分频域滤波器,有效建模语音中的长程依赖关系。同时,专门设计了动态滤波策略与稀疏正则化方法,以增强全局滤波器的性能并防止过拟合。进一步,我们构建了双流TDNN(DS-TDNN),通过拆分基础通道降低复杂度,并采用全局滤波器提升识别性能。在Voxceleb与SITW数据库上的实验表明:与ECAPA-TDNN相比,DS-TDNN在实现约10%性能提升的同时,计算复杂度与参数量分别降低28%以上和15%以上。此外,在处理长时语音时,相较于其他主流基线系统,DS-TDNN在效率与有效性之间取得了最优平衡。最终,可视化分析与详细消融研究进一步揭示了DS-TDNN的优势。