As Speech Large Language Models (Speech LLMs) become increasingly integrated into voice-based applications, ensuring their robustness against manipulative or adversarial input becomes critical. Although prior work has studied adversarial attacks in text-based LLMs and vision-language models, the unique cognitive and perceptual challenges of speech-based interaction remain underexplored. In contrast, speech presents inherent ambiguity, continuity, and perceptual diversity, which make adversarial attacks more difficult to detect. In this paper, we introduce gaslighting attacks, strategically crafted prompts designed to mislead, override, or distort model reasoning as a means to evaluate the vulnerability of Speech LLMs. Specifically, we construct five manipulation strategies: Anger, Cognitive Disruption, Sarcasm, Implicit, and Professional Negation, designed to test model robustness across varied tasks. It is worth noting that our framework captures both performance degradation and behavioral responses, including unsolicited apologies and refusals, to diagnose different dimensions of susceptibility. Moreover, acoustic perturbation experiments are conducted to assess multi-modal robustness. To quantify model vulnerability, comprehensive evaluation across 5 Speech and multi-modal LLMs on over 10,000 test samples from 5 diverse datasets reveals an average accuracy drop of 24.3% under the five gaslighting attacks, indicating significant behavioral vulnerability. These findings highlight the need for more resilient and trustworthy speech-based AI systems.
翻译:随着语音大语言模型(Speech LLMs)日益融入各类语音应用,确保其对操纵性或对抗性输入的鲁棒性变得至关重要。尽管先前研究已涉及文本大语言模型和视觉语言模型中的对抗性攻击,但语音交互所特有的认知与感知挑战仍未得到充分探索。相比之下,语音呈现出固有的模糊性、连续性和感知多样性,这使得对抗性攻击更难以被检测。本文提出了煤气灯攻击——一种精心设计的提示策略,旨在误导、覆盖或扭曲模型推理,以评估语音大语言模型的脆弱性。具体而言,我们构建了五种操纵策略:愤怒、认知扰乱、讽刺、隐晦和职业否定,用于测试模型在不同任务上的鲁棒性。值得注意的是,我们的框架同时捕捉了性能下降和行为响应(包括主动道歉和拒绝回答),以诊断不同维度的易感性。此外,还进行了声学扰动实验以评估多模态鲁棒性。为量化模型脆弱性,我们在5个语音及多模态大语言模型上,基于来自5个不同数据集的超过10,000个测试样本进行了全面评估,结果显示在五种煤气灯攻击下平均准确率下降24.3%,表明存在显著的行为脆弱性。这些发现凸显了构建更具韧性和可信赖的语音人工智能系统的迫切需求。