Previous studies demonstrate the impressive performance of residual neural networks (ResNet) in speaker verification. The ResNet models treat the time and frequency dimensions equally. They follow the default stride configuration designed for image recognition, where the horizontal and vertical axes exhibit similarities. This approach ignores the fact that time and frequency are asymmetric in speech representation. In this paper, we address this issue and look for optimal stride configurations specifically tailored for speaker verification. We represent the stride space on a trellis diagram, and conduct a systematic study on the impact of temporal and frequency resolutions on the performance and further identify two optimal points, namely Golden Gemini, which serves as a guiding principle for designing 2D ResNet-based speaker verification models. By following the principle, a state-of-the-art ResNet baseline model gains a significant performance improvement on VoxCeleb, SITW, and CNCeleb datasets with 7.70%/11.76% average EER/minDCF reductions, respectively, across different network depths (ResNet18, 34, 50, and 101), while reducing the number of parameters by 16.5% and FLOPs by 4.1%. We refer to it as Gemini ResNet. Further investigation reveals the efficacy of the proposed Golden Gemini operating points across various training conditions and architectures. Furthermore, we present a new benchmark, namely the Gemini DF-ResNet, using a cutting-edge model.
翻译:以往研究表明,残差神经网络(ResNet)在说话人验证中展现了卓越性能。然而,ResNet模型平等对待时间和频率维度,遵循为图像识别设计的默认步长配置(其中水平轴与垂直轴具有相似性)。这种方法忽略了语音表征中时间与频率的非对称特性。本文针对该问题,探索专门为说话人验证优化的最佳步长配置。我们在网格图上表征步长空间,系统研究时域与频域分辨率对性能的影响,进而识别出两个最优操作点,即"黄金双子座"(Golden Gemini),该原则可作为设计基于二维ResNet的说话人验证模型的指导准则。遵循该原则,最先进的ResNet基线模型在VoxCeleb、SITW和CNCeleb数据集上分别取得平均7.70%/11.76%的等错误率(EER)/最小检测代价函数(minDCF)显著降低,同时在不同网络深度(ResNet18、34、50和101)下减少16.5%的参数量和4.1%的浮点运算数(FLOPs)。我们将其称为Gemini ResNet。进一步研究表明,所提出的黄金双子座操作点在不同训练条件和架构下均具有效性。此外,我们采用前沿模型,构建了名为Gemini DF-ResNet的新基准。