This paper reveals the phenomenon of parameter heterogeneity in large language models (LLMs). We find that a small subset of ``cherry'' parameters exhibit a disproportionately large influence on model performance, while the vast majority of parameters have minimal impact. This heterogeneity is found to be prevalent across different model families, scales, and types. Motivated by this observation, we propose CherryQ, a novel quantization method that unifies the optimization of mixed-precision parameters. CherryQ identifies and preserves the critical cherry parameters in high precision while aggressively quantizing the remaining parameters to low precision. Extensive experiments demonstrate the effectiveness of CherryQ. CherryQ outperforms existing quantization approaches in terms of perplexity and downstream task performance. Notably, our 3-bit quantized Vicuna-1.5 exhibits competitive performance compared to their 16-bit counterparts. These findings highlight the potential of CherryQ for enabling efficient deployment of LLMs by taking advantage of parameter heterogeneity.
翻译:本文揭示了大型语言模型(LLMs)中的参数异质性现象。研究发现,少量"重点"参数对模型性能产生不成比例的巨大影响,而绝大多数参数的影响微乎其微。这种异质性普遍存在于不同模型家族、规模及类型中。基于此观察,我们提出CherryQ——一种统一优化混合精度参数的新型量化方法。CherryQ能够识别并保留关键的重点参数采用高精度表示,同时将剩余参数激进量化为低精度。大量实验验证了CherryQ的有效性。在困惑度与下游任务性能方面,CherryQ均优于现有量化方法。值得注意的是,采用3比特量化后的Vicuna-1.5模型展现出与16比特版本相媲美的性能。这些发现凸显了CherryQ通过利用参数异质性实现LLMs高效部署的潜力。