With the huge technological advances introduced by deep learning in audio & speech processing, many novel synthetic speech techniques achieved incredible realistic results. As these methods generate realistic fake human voices, they can be used in malicious acts such as people imitation, fake news, spreading, spoofing, media manipulations, etc. Hence, the ability to detect synthetic or natural speech has become an urgent necessity. Moreover, being able to tell which algorithm has been used to generate a synthetic speech track can be of preeminent importance to track down the culprit. In this paper, a novel strategy is proposed to attribute a synthetic speech track to the generator that is used to synthesize it. The proposed detector transforms the audio into log-mel spectrogram, extracts features using CNN, and classifies it between five known and unknown algorithms, utilizing semi-supervision and ensemble to improve its robustness and generalizability significantly. The proposed detector is validated on two evaluation datasets consisting of a total of 18,000 weakly perturbed (Eval 1) & 10,000 strongly perturbed (Eval 2) synthetic speeches. The proposed method outperforms other top teams in accuracy by 12-13% on Eval 2 and 1-2% on Eval 1, in the IEEE SP Cup challenge at ICASSP 2022.
翻译:随着深度学习在音频与语音处理领域的巨大技术进步,许多新型合成语音技术取得了令人难以置信的逼真效果。由于这些方法能够生成逼真的虚假人声,它们可能被用于恶意行为,如模仿他人、传播假新闻、欺骗、媒体操纵等。因此,检测合成语音或自然语音的能力已成为迫切需求。此外,能够识别出用于生成合成语音轨道的算法,对于追查肇事者至关重要。本文提出了一种新策略,用于将合成语音轨道归因于生成它的合成器。所提出的检测器将音频转换为对数梅尔声谱图,利用卷积神经网络提取特征,并在五个已知和未知算法之间进行分类,通过半监督学习与集成方法显著提升其鲁棒性和泛化能力。该检测器在两个评估数据集上进行了验证,分别包含18000个弱扰动(评估集1)和10000个强扰动(评估集2)合成语音样本。在ICASSP 2022的IEEE SP Cup挑战赛中,所提出的方法在评估集2上的准确率比其他顶尖团队高出12-13%,在评估集1上高出1-2%。