The wav2vec 2.0 and integrated spectro-temporal graph attention network (AASIST) based countermeasure achieves great performance in speech anti-spoofing. However, current spoof speech detection systems have fixed training and evaluation durations, while the performance degrades significantly during short utterance evaluation. To solve this problem, AASIST can be improved to AASIST2 by modifying the residual blocks to Res2Net blocks. The modified Res2Net blocks can extract multi-scale features and improve the detection performance for speech of different durations, thus improving the short utterance evaluation performance. On the other hand, adaptive large margin fine-tuning (ALMFT) has achieved performance improvement in short utterance speaker verification. Therefore, we apply Dynamic Chunk Size (DCS) and ALMFT training strategies in speech anti-spoofing to further improve the performance of short utterance evaluation. Experiments demonstrate that the proposed AASIST2 improves the performance of short utterance evaluation while maintaining the performance of regular evaluation on different datasets.
翻译:基于wav2vec 2.0与谱-时图注意力网络(AASIST)的对抗措施在语音防欺骗中取得了优异性能。然而,当前欺骗语音检测系统的训练与评估时长固定,在短语音评估场景下性能显著下降。为应对该问题,可将AASIST改进为AASIST2,其通过将残差块替换为Res2Net块实现。改进后的Res2Net块能提取多尺度特征,提升对不同时长语音的检测性能,进而改善短语音评估效果。此外,自适应大间隔微调(ALMFT)已在短语音说话人验证中取得性能提升。因此,我们将动态块切分(DCS)与ALMFT训练策略引入语音防欺骗任务,以进一步提升短语音评估性能。实验表明,所提出的AASIST2在保持常规评估性能的同时,提升了不同数据集上短语音评估的效果。