Benchmarking initiatives support the meaningful comparison of competing solutions to prominent problems in speech and language processing. Successive benchmarking evaluations typically reflect a progressive evolution from ideal lab conditions towards to those encountered in the wild. ASVspoof, the spoofing and deepfake detection initiative and challenge series, has followed the same trend. This article provides a summary of the ASVspoof 2021 challenge and the results of 54 participating teams that submitted to the evaluation phase. For the logical access (LA) task, results indicate that countermeasures are robust to newly introduced encoding and transmission effects. Results for the physical access (PA) task indicate the potential to detect replay attacks in real, as opposed to simulated physical spaces, but a lack of robustness to variations between simulated and real acoustic environments. The Deepfake (DF) task, new to the 2021 edition, targets solutions to the detection of manipulated, compressed speech data posted online. While detection solutions offer some resilience to compression effects, they lack generalization across different source datasets. In addition to a summary of the top-performing systems for each task, new analyses of influential data factors and results for hidden data subsets, the article includes a review of post-challenge results, an outline of the principal challenge limitations and a road-map for the future of ASVspoof.
翻译:基准测试倡议支持对语音与语言处理领域中重要问题的竞争性解决方案进行有意义的比较。连续的基准测试评估通常反映了从理想实验室条件向现实场景条件逐步演进的趋势。ASVspoof(欺骗与深度伪造检测倡议及挑战系列)也遵循了这一趋势。本文总结了ASVspoof 2021挑战赛及提交至评估阶段的54支参赛队伍的结果。对于逻辑访问(LA)任务,结果表明反制措施对新引入的编码和传输效应具有鲁棒性。物理访问(PA)任务的结果表明,在真实物理空间(而非模拟空间)中检测重放攻击具有潜力,但对模拟与真实声学环境之间的变化缺乏鲁棒性。深度伪造(DF)任务是2021年新增项目,旨在针对在线发布的经过篡改和压缩的语音数据提供检测方案。尽管检测方案对压缩效应具有一定的抗性,但它们在不同源数据集之间缺乏泛化能力。除了总结每项任务的顶级系统性能外,本文还对影响数据因素和隐藏数据子集的结果进行了新分析,并回顾了挑战赛后结果、概述了主要挑战局限性以及ASVspoof的未来发展路线图。