Recently, Transformer-based architectures have been explored for speaker embedding extraction. Although the Transformer employs the self-attention mechanism to efficiently model the global interaction between token embeddings, it is inadequate for capturing short-range local context, which is essential for the accurate extraction of speaker information. In this study, we enhance the Transformer with the locality modeling in two directions. First, we propose the Locality-Enhanced Conformer (LE-Confomer) by introducing depth-wise convolution and channel-wise attention into the Conformer blocks. Second, we present the Speaker Swin Transformer (SST) by adapting the Swin Transformer, originally proposed for vision tasks, into speaker embedding network. We evaluate the proposed approaches on the VoxCeleb datasets and a large-scale Microsoft internal multilingual (MS-internal) dataset. The proposed models achieve 0.75% EER on VoxCeleb 1 test set, outperforming the previously proposed Transformer-based models and CNN-based models, such as ResNet34 and ECAPA-TDNN. When trained on the MS-internal dataset, the proposed models achieve promising results with 14.6% relative reduction in EER over the Res2Net50 model.
翻译:近期,基于Transformer的架构已被探索用于说话人嵌入提取。尽管Transformer利用自注意力机制有效建模词元嵌入之间的全局交互,但其在捕获短程局部上下文方面存在不足,而这对于准确提取说话人信息至关重要。本研究从两个方向增强Transformer的局部性建模能力。首先,我们提出局部性增强Conformer(简称LE-Confomer),通过在Conformer模块中引入深度可分离卷积和通道注意力机制。其次,我们提出说话人Swin Transformer(简称SST),将最初面向视觉任务的Swin Transformer适配为说话人嵌入网络。我们在VoxCeleb数据集和大规模微软内部多语言(MS-internal)数据集上评估所提方法。所提模型在VoxCeleb 1测试集上达到0.75%的等错误率(EER),优于先前提出的基于Transformer的模型和基于CNN的模型(如ResNet34和ECAPA-TDNN)。在MS-internal数据集上训练时,所提模型相较Res2Net50模型实现EER相对降低14.6%的显著效果。