The rise of singing voice synthesis presents critical challenges to artists and industry stakeholders over unauthorized voice usage. Unlike synthesized speech, synthesized singing voices are typically released in songs containing strong background music that may hide synthesis artifacts. Additionally, singing voices present different acoustic and linguistic characteristics from speech utterances. These unique properties make singing voice deepfake detection a relevant but significantly different problem from synthetic speech detection. In this work, we propose the singing voice deepfake detection task. We first present SingFake, the first curated in-the-wild dataset consisting of 28.93 hours of bonafide and 29.40 hours of deepfake song clips in five languages from 40 singers. We provide a train/validation/test split where the test sets include various scenarios. We then use SingFake to evaluate four state-of-the-art speech countermeasure systems trained on speech utterances. We find these systems lag significantly behind their performance on speech test data. When trained on SingFake, either using separated vocal tracks or song mixtures, these systems show substantial improvement. However, our evaluations also identify challenges associated with unseen singers, communication codecs, languages, and musical contexts, calling for dedicated research into singing voice deepfake detection. The SingFake dataset and related resources are available at https://www.singfake.org/.
翻译:歌唱语音合成的兴起对艺术家和行业利益相关者带来了未经授权使用声音的严峻挑战。与合成语音不同,合成歌唱语音通常出现在包含强烈背景音乐的歌曲中,这可能会掩盖合成痕迹。此外,歌唱语音具有与语音话语不同的声学和语言学特征。这些独特性质使得歌唱语音深度伪造检测成为一个相关但与合成语音检测截然不同的难题。本文中,我们提出了歌唱语音深度伪造检测任务。首先介绍了SingFake,这是首个精心策划的真实场景数据集,包含来自40位歌手的28.93小时真实和29.40小时深度伪造歌曲片段,涵盖五种语言。我们提供了训练/验证/测试划分,其中测试集包含多种场景。随后,我们利用SingFake评估了四种基于语音话语训练的先进语音反欺骗系统。结果发现,这些系统在歌唱语音测试数据上的表现远逊于其在语音测试数据上的性能。当使用分离的人声轨道或歌曲混合训练时,这些系统在SingFake上取得了显著改进。然而,我们的评估也揭示了与未见歌手、通信编解码器、语言和音乐背景相关的挑战,呼吁对歌唱语音深度伪造检测开展专项研究。SingFake数据集及相关资源可在https://www.singfake.org/获取。