Deep Learning (DL) models have been popular nowadays to execute different speech-related tasks, including automatic speech recognition (ASR). As ASR is being used in different real-time scenarios, it is important that the ASR model remains efficient against minor perturbations to the input. Hence, evaluating efficiency robustness of the ASR model is the need of the hour. We show that popular ASR models like Speech2Text model and Whisper model have dynamic computation based on different inputs, causing dynamic efficiency. In this work, we propose SlothSpeech, a denial-of-service attack against ASR models, which exploits the dynamic behaviour of the model. SlothSpeech uses the probability distribution of the output text tokens to generate perturbations to the audio such that efficiency of the ASR model is decreased. We find that SlothSpeech generated inputs can increase the latency up to 40X times the latency induced by benign input.
翻译:深度学习模型目前已广泛用于执行各种语音相关任务,包括自动语音识别。由于自动语音识别正被应用于多种实时场景,确保ASR模型对输入中的微小扰动保持高效性至关重要。因此,评估ASR模型的效率鲁棒性是当务之急。我们证明,像Speech2Text和Whisper这类主流ASR模型会根据不同输入动态调整计算量,从而导致效率的动态变化。本文提出SlothSpeech——一种针对ASR模型的拒绝服务攻击方法,该攻击利用模型的动态行为特性。SlothSpeech通过输出文本令牌的概率分布生成针对音频的扰动,从而降低ASR模型的效率。我们发现,相较良性输入,SlothSpeech生成的输入可将延迟提升高达40倍。