Due to the successful application of deep learning, audio spoofing detection has made significant progress. Spoofed audio with speech synthesis or voice conversion can be well detected by many countermeasures. However, an automatic speaker verification system is still vulnerable to spoofing attacks such as replay or Deep-Fake audio. Deep-Fake audio means that the spoofed utterances are generated using text-to-speech (TTS) and voice conversion (VC) algorithms. Here, we propose a novel framework based on hybrid features with the self-attention mechanism. It is expected that hybrid features can be used to get more discrimination capacity. Firstly, instead of only one type of conventional feature, deep learning features and Mel-spectrogram features will be extracted by two parallel paths: convolution neural networks and a short-time Fourier transform (STFT) followed by Mel-frequency. Secondly, features will be concatenated by a max-pooling layer. Thirdly, there is a Self-attention mechanism for focusing on essential elements. Finally, ResNet and a linear layer are built to get the results. Experimental results reveal that the hybrid features, compared with conventional features, can cover more details of an utterance. We achieve the best Equal Error Rate (EER) of 9.67\% in the physical access (PA) scenario and 8.94\% in the Deep fake task on the ASVspoof 2021 dataset. Compared with the best baseline system, the proposed approach improves by 74.60\% and 60.05\%, respectively.
翻译:由于深度学习的成功应用,音频欺骗检测取得了显著进展。通过语音合成或语音转换生成的欺骗音频能被多种反制措施有效检测。然而,自动说话人验证系统仍易受重放或深度伪造音频等欺骗攻击的影响。深度伪造音频指利用文本转语音(TTS)和语音转换(VC)算法生成的欺骗语句。本文提出了一种基于混合特征与自注意力机制的新型框架,旨在利用混合特征获得更强的辨别能力。首先,区别于仅使用单一传统特征,深度学习特征与梅尔频谱特征将通过两条并行路径提取:卷积神经网络与短时傅里叶变换(STFT)后接梅尔频率处理。其次,特征通过最大池化层进行拼接。随后引入自注意力机制聚焦关键要素。最后,构建残差网络(ResNet)与线性层输出结果。实验结果表明,与传统特征相比,混合特征能覆盖语句中更多细节信息。在ASVspoof 2021数据集上,本方法在物理访问(PA)场景下取得9.67%的最佳等错误率(EER),在深度伪造任务中取得8.94%的EER。与最优基线系统相比,所提方法分别提升了74.60%和60.05%。