Audio adversarial examples (AEs) have posed significant security challenges to real-world speaker recognition systems. Most black-box attacks still require certain information from the speaker recognition model to be effective (e.g., keeping probing and requiring the knowledge of similarity scores). This work aims to push the practicality of the black-box attacks by minimizing the attacker's knowledge about a target speaker recognition model. Although it is not feasible for an attacker to succeed with completely zero knowledge, we assume that the attacker only knows a short (or a few seconds) speech sample of a target speaker. Without any probing to gain further knowledge about the target model, we propose a new mechanism, called parrot training, to generate AEs against the target model. Motivated by recent advancements in voice conversion (VC), we propose to use the one short sentence knowledge to generate more synthetic speech samples that sound like the target speaker, called parrot speech. Then, we use these parrot speech samples to train a parrot-trained(PT) surrogate model for the attacker. Under a joint transferability and perception framework, we investigate different ways to generate AEs on the PT model (called PT-AEs) to ensure the PT-AEs can be generated with high transferability to a black-box target model with good human perceptual quality. Real-world experiments show that the resultant PT-AEs achieve the attack success rates of 45.8% - 80.8% against the open-source models in the digital-line scenario and 47.9% - 58.3% against smart devices, including Apple HomePod (Siri), Amazon Echo, and Google Home, in the over-the-air scenario.
翻译:音频对抗样本(AEs)已对现实世界的说话人识别系统构成重大安全挑战。大多数黑盒攻击仍需从说话人识别模型中获取特定信息方能有效(例如,持续探测并需要相似度分数等知识)。本研究旨在通过最小化攻击者对目标说话人识别模型的已知信息,推动黑盒攻击的实用性。虽然攻击者完全零知识的情况下不可能成功,但我们假设攻击者仅掌握目标说话人的一段简短(或数秒)语音样本。在不进行任何探测以获取目标模型进一步知识的前提下,我们提出一种名为"鹦鹉训练"的新机制来生成针对目标模型的对抗样本。受语音转换(VC)领域最新进展启发,我们提出利用单句语音知识生成更多听起来像目标说话人的合成语音样本,称为"鹦鹉语音"。随后,我们使用这些鹦鹉语音样本来训练一个攻击者专用的鹦鹉训练(PT)代理模型。在联合可迁移性与感知性框架下,我们研究了在PT模型上生成对抗样本(称为PT-AEs)的不同方法,以确保PT-AEs能够在保持良好人类感知质量的同时,具备对黑盒目标模型的高可迁移性。真实世界实验表明,在数字线路场景下,生成的PT-AEs对开源模型的攻击成功率达45.8%-80.8%;在无线传输场景下,对包括Apple HomePod(Siri)、Amazon Echo和Google Home在内的智能设备的攻击成功率达47.9%-58.3%。