Speaker recognition technology is applied in various tasks ranging from personal virtual assistants to secure access systems. However, the robustness of these systems against adversarial attacks, particularly to additive perturbations, remains a significant challenge. In this paper, we pioneer applying robustness certification techniques to speaker recognition, originally developed for the image domain. In our work, we cover this gap by transferring and improving randomized smoothing certification techniques against norm-bounded additive perturbations for classification and few-shot learning tasks to speaker recognition. We demonstrate the effectiveness of these methods on VoxCeleb 1 and 2 datasets for several models. We expect this work to improve voice-biometry robustness, establish a new certification benchmark, and accelerate research of certification methods in the audio domain.
翻译:说话人识别技术广泛应用于从个人虚拟助手到安全接入系统的多种任务中。然而,这些系统对抗对抗攻击(尤其是加性扰动)的鲁棒性仍是一项重大挑战。本文首次将最初为图像领域开发的鲁棒性认证技术应用于说话人识别。我们通过迁移并改进针对范数约束加性扰动的随机平滑认证技术,填补了这一空白,使其适用于说话人识别的分类和少样本学习任务。我们在VoxCeleb 1和VoxCeleb 2数据集上验证了这些方法在多个模型上的有效性。期待本工作能提升语音生物特征的鲁棒性,建立新的认证基准,并加速音频领域认证方法的研究进程。