The malicious use of deep speech synthesis models may pose significant threat to society. Therefore, many studies have emerged to detect the so-called ``deepfake audio". However, these studies focus on the binary detection of real audio and fake audio. For some realistic application scenarios, it is needed to know what tool or model generated the deepfake audio. This raises a question: Can we recognize the system fingerprints of deepfake audio? Therefore, in this paper, we propose a deepfake audio dataset for system fingerprint recognition (SFR) and conduct an initial investigation. We collected the dataset from five speech synthesis systems using the latest state-of-the-art deep learning technologies, including both clean and compressed sets. In addition, to facilitate the further development of system fingerprint recognition methods, we give researchers some benchmarks that can be compared, and research findings. The dataset will be publicly available.
翻译:恶意使用深度语音合成模型可能对社会构成重大威胁。为此,已有大量研究致力于检测所谓的"深度伪造音频"。然而,现有研究主要聚焦于真实音频与伪造音频的二元检测。在部分实际应用场景中,我们需要知晓深度伪造音频是由何种工具或模型生成的。由此引发一个问题:我们能否识别深度伪造音频的系统指纹?针对这一问题,本文提出了一套用于系统指纹识别(SFR)的深度伪造音频数据集并展开初步探究。该数据集采集自采用最新先进深度学习技术的五个语音合成系统,涵盖纯净音频与压缩音频两种形式。此外,为促进系统指纹识别方法的进一步发展,我们为研究人员提供了可供对比的基准方法及研究成果。该数据集将公开提供。