Audio adversarial examples (AEs) have posed significant security challenges to real-world speaker recognition systems. Most black-box attacks still require certain information from the speaker recognition model to be effective (e.g., keeping probing and requiring the knowledge of similarity scores). This work aims to push the practicality of the black-box attacks by minimizing the attacker's knowledge about a target speaker recognition model. Although it is not feasible for an attacker to succeed with completely zero knowledge, we assume that the attacker only knows a short (or a few seconds) speech sample of a target speaker. Without any probing to gain further knowledge about the target model, we propose a new mechanism, called parrot training, to generate AEs against the target model. Motivated by recent advancements in voice conversion (VC), we propose to use the one short sentence knowledge to generate more synthetic speech samples that sound like the target speaker, called parrot speech. Then, we use these parrot speech samples to train a parrot-trained(PT) surrogate model for the attacker. Under a joint transferability and perception framework, we investigate different ways to generate AEs on the PT model (called PT-AEs) to ensure the PT-AEs can be generated with high transferability to a black-box target model with good human perceptual quality. Real-world experiments show that the resultant PT-AEs achieve the attack success rates of 45.8% - 80.8% against the open-source models in the digital-line scenario and 47.9% - 58.3% against smart devices, including Apple HomePod (Siri), Amazon Echo, and Google Home, in the over-the-air scenario.
翻译:音频对抗样本(AEs)对现实世界中的说话人识别系统构成了重大安全挑战。大多数黑盒攻击仍需从说话人识别模型中获取特定信息才能生效(例如持续探测并获取相似度分数)。本研究旨在通过最小化攻击者对目标说话人识别模型的先验知识,提升黑盒攻击的实际可用性。尽管攻击者在完全零知识条件下无法成功,我们假设攻击者仅拥有目标说话人的一段短时(数秒)语音样本。在不进行任何探测以获取目标模型额外知识的前提下,我们提出名为"鹦鹉训练"的新机制,用于生成针对目标模型的对抗样本。受语音转换(VC)技术最新进展的启发,我们提出利用单句语音知识合成更多类似目标说话人声音的语音样本,称为鹦鹉语音。随后,使用这些鹦鹉语音样本为攻击者训练鹦鹉训练(PT)替代模型。基于联合可迁移性与感知质量框架,我们探索了在PT模型上生成对抗样本(称为PT-AEs)的不同方法,以确保生成的PT-AEs既能高可迁移性地攻击黑盒目标模型,又能保持良好的听觉感知质量。现实世界实验表明,最终获得的PT-AEs在数字线路场景下对开源模型的攻击成功率达45.8%-80.8%,在无线传输场景下对苹果HomePod(Siri)、亚马逊Echo和谷歌Home等智能设备的攻击成功率达47.9%-58.3%。