Audio deepfake detection is an emerging topic in the artificial intelligence community. The second Audio Deepfake Detection Challenge (ADD 2023) aims to spur researchers around the world to build new innovative technologies that can further accelerate and foster research on detecting and analyzing deepfake speech utterances. Different from previous challenges (e.g. ADD 2022), ADD 2023 focuses on surpassing the constraints of binary real/fake classification, and actually localizing the manipulated intervals in a partially fake speech as well as pinpointing the source responsible for generating any fake audio. Furthermore, ADD 2023 includes more rounds of evaluation for the fake audio game sub-challenge. The ADD 2023 challenge includes three subchallenges: audio fake game (FG), manipulation region location (RL) and deepfake algorithm recognition (AR). This paper describes the datasets, evaluation metrics, and protocols. Some findings are also reported in audio deepfake detection tasks.
翻译:音频深度伪造检测是人工智能领域的一个新兴课题。第二届音频深度伪造检测挑战赛(ADD 2023)旨在激励全球研究人员开发创新技术,以进一步加速和促进深度伪造语音片段的检测与分析研究。与以往挑战(如ADD 2022)不同,ADD 2023聚焦于突破二元真/假分类的限制,致力于定位部分伪造语音中的篡改区间,并精准识别生成任何伪造音频的来源。此外,ADD 2023在伪造音频游戏子挑战中增设了更多评估轮次。本次挑战包含三个子任务:音频伪造游戏(FG)、篡改区域定位(RL)及深度伪造算法识别(AR)。本文描述了数据集、评估指标及协议,并报告了音频深度伪造检测任务中的若干研究发现。