In this paper, we propose the multi-perspective information fusion (MPIF) Res2Net with random Specmix for fake speech detection (FSD). The main purpose of this system is to improve the model's ability to learn precise forgery information for FSD task in low-quality scenarios. The task of random Specmix, a data augmentation, is to improve the generalization ability of the model and enhance the model's ability to locate discriminative information. Specmix cuts and pastes the frequency dimension information of the spectrogram in the same batch of samples without introducing other data, which helps the model to locate the really useful information. At the same time, we randomly select samples for augmentation to reduce the impact of data augmentation directly changing all the data. Once the purpose of helping the model to locate information is achieved, it is also important to reduce unnecessary information. The role of MPIF-Res2Net is to reduce redundant interference information. Deceptive information from a single perspective is always similar, so the model learning this similar information will produce redundant spoofing clues and interfere with truly discriminative information. The proposed MPIF-Res2Net fuses information from different perspectives, making the information learned by the model more diverse, thereby reducing the redundancy caused by similar information and avoiding interference with the learning of discriminative information. The results on the ASVspoof 2021 LA dataset demonstrate the effectiveness of our proposed method, achieving EER and min-tDCF of 3.29% and 0.2557, respectively.
翻译:本文提出了一种结合随机Specmix的多视角信息融合(MPIF)Res2Net方法,用于虚假语音检测(FSD)。该系统的主要目的是提升模型在低质量场景下学习精确伪造信息的能力。随机Specmix作为一种数据增强技术,其任务是提高模型的泛化能力,并增强其定位判别性信息的能力。Specmix通过在同一批样本中剪切并粘贴频谱图的频率维度信息(无需引入额外数据),帮助模型定位真正有用的信息。同时,我们随机选择样本进行增强,以减少数据增强直接改变所有数据带来的影响。在帮助模型定位信息的目标达成后,减少不必要的信息同样重要。MPIF-Res2Net的作用正在于降低冗余干扰信息。单一视角的欺骗性信息往往具有相似性,因此模型学习此类相似信息会生成冗余的欺骗线索,干扰真实判别性信息的学习。本文提出的MPIF-Res2Net融合了不同视角的信息,使模型学习到的信息更加多样,从而降低相似信息导致的冗余,并避免对判别性信息学习的干扰。在ASVspoof 2021 LA数据集上的实验结果验证了所提方法的有效性,等错误率(EER)和最小检测代价函数(min-tDCF)分别达到3.29%和0.2557。