LLM quantization has become essential for memory-efficient deployment. Recent work has shown that quantization schemes can pose critical security risks: an adversary may release a model that appears benign in full precision but exhibits malicious behavior once quantized by users. However, existing quantization-conditioned attacks have been limited to relatively simple quantization methods, where the attacker can estimate weight regions that remain invariant under the target quantization. Notably, prior attacks have consistently failed to compromise more popular and sophisticated schemes, limiting their practical impact. In this work, we introduce the first quantization-conditioned attack that consistently induces malicious behavior that can be triggered by a broad range of advanced quantization techniques, including AWQ, GPTQ, and GGUF I-quants. Our attack exploits a simple property shared by many modern quantization methods: large outliers can cause other weights to be rounded to zero. Consequently, by injecting outliers into specific weight blocks, an adversary can therefore induce a targeted, predictable weight collapse in the model. This effect can be used to craft seemingly benign full-precision models that exhibit a wide range of malicious behaviors after quantization. Through extensive evaluation across three attack scenarios and LLMs, we show that our attack achieves high success rates against a broad range of quantization methods on which prior attacks fail. Our results demonstrate, for the first time, that the security risks of quantization are not restricted to simpler schemes but are broadly relevant across complex, widely-used quantization methods.
翻译:LLM量化已成为内存高效部署的关键技术。近期研究表明,量化方案可能带来严重安全风险:攻击者可发布看似全精度无害的模型,经用户量化后却展现出恶意行为。然而,现有面向量化条件的攻击仅局限于相对简单的量化方法,其前提是攻击者能够估计在目标量化下保持不变的权重区域。值得注意的是,先前的攻击始终无法攻破更流行且复杂的量化方案,这限制了其实际影响。本文首次提出一种面向量化条件的攻击方法,能够使模型在现代高级量化技术(包括AWQ、GPTQ和GGUF I-quants)的一致触发下产生恶意行为。我们的攻击利用了现代量化方法共有的简单特性:较大的异常值会导致其他权重被舍入为零。因此,通过向特定权重块注入异常值,攻击者可诱导模型产生定向且可预测的权重坍塌效应。该效应可用于构建看似正常的全精度模型,这些模型在量化后能够展现多种恶意行为。通过在三种攻击场景和多个LLM上的广泛评估,我们证明该攻击方法对先前攻击失效的多种量化方案均具有高成功率。我们的结果首次表明,量化的安全风险不仅局限于简单方案,而是广泛存在于复杂且广泛使用的量化方法中。