Tools to generate high quality synthetic speech signal that is perceptually indistinguishable from speech recorded from human speakers are easily available. Several approaches have been proposed for detecting synthetic speech. Many of these approaches use deep learning methods as a black box without providing reasoning for the decisions they make. This limits the interpretability of these approaches. In this paper, we propose Disentangled Spectrogram Variational Auto Encoder (DSVAE) which is a two staged trained variational autoencoder that processes spectrograms of speech using disentangled representation learning to generate interpretable representations of a speech signal for detecting synthetic speech. DSVAE also creates an activation map to highlight the spectrogram regions that discriminate synthetic and bona fide human speech signals. We evaluated the representations obtained from DSVAE using the ASVspoof2019 dataset. Our experimental results show high accuracy (>98%) on detecting synthetic speech from 6 known and 10 out of 11 unknown speech synthesizers. We also visualize the representation obtained from DSVAE for 17 different speech synthesizers and verify that they are indeed interpretable and discriminate bona fide and synthetic speech from each of the synthesizers.
翻译:能够生成与人类录音在感知上无法区分的高质量合成语音信号的工具已广泛可得。针对合成语音检测,已有多种方法被提出。其中许多方法将深度学习模型视为黑箱,未对其决策提供解释,这限制了这些方法的可解释性。本文提出解耦语谱图变分自编码器(DSVAE),这是一种两阶段训练的变分自编码器,通过解耦表征学习处理语音语谱图,生成用于检测合成语音的可解释语音信号表征。DSVAE还构建激活图谱以突出显示区分合成语音与真实人类语音信号的语谱区域。我们使用ASVspoof2019数据集对DSVAE获取的表征进行评估。实验结果表明,该方法对来自6种已知及11种未知语音合成器中的10种所生成的合成语音检测准确率超过98%。我们还可视化了DSVAE对17种不同语音合成器所获取的表征,并验证这些表征确实具有可解释性,能够有效区分真实语音与各合成器生成的合成语音。