We present \textbf{ITQ3\_S} (Interleaved Ternary Quantization -- Specialized), a novel 3-bit weight quantization format for large language models (LLMs) that integrates \textbf{TurboQuant (TQ)}, a rotation-domain adaptive quantization strategy based on the Fast Walsh-Hadamard Transform (FWHT). Conventional 3-bit quantization methods suffer from catastrophic precision loss caused by heavy-tailed weight distributions and inter-channel outliers. ITQ3\_S addresses this fundamental limitation by pre-rotating the weight space via FWHT prior to quantization, effectively spreading outlier energy across the entire vector and inducing a near-Gaussian distribution amenable to uniform ternary coding. Critically, we derive a mathematically rigorous dequantization procedure that inverts the FWHT exactly using a 256-point Inverse Walsh-Hadamard Transform fused into the CUDA shared-memory loading stage, ensuring zero-error round-trip fidelity between offline quantization and online inference. We prove that for any weight vector $\mathbf{w} \in \mathbb{R}^{256}$ processed by our pipeline, the reconstruction satisfies $\|\hat{\mathbf{w}} - \mathbf{w}\|_2 \leq ε_q$, where $ε_q$ is determined solely by the ternary quantization grid and is strictly smaller than any uniform 3-bit baseline under equal bit-budget constraints. Empirically, on the NVIDIA RTX 5090 (Blackwell architecture), ITQ3\_S achieves perplexity competitive with FP16 baselines while delivering throughput exceeding 1.5$\times$ that of 4-bit alternatives, owing to optimized DP4A and Tensor Core scheduling in the interleaved memory layout. Our results establish ITQ3\_S as a practical, mathematically grounded solution for high-fidelity LLM deployment on consumer-grade hardware.


翻译:我们提出\textbf{ITQ3\_S}(交错三值量化——专用型),这是一种面向大语言模型的新型3位权重量化格式,其融合了基于快速沃尔什-阿达玛变换(FWHT)的旋转域自适应量化策略\textbf{TurboQuant(TQ)}。传统3位量化方法因重尾权值分布与通道间异常值而遭受灾难性精度损失。ITQ3\_S通过在量化前利用FWHT预旋转权值空间突破这一根本限制,有效将异常值能量分散至整个向量,诱导出适于均匀三值编码的近高斯分布。关键之处在于,我们推导了严格数学化的反量化流程——通过将融合于CUDA共享内存加载阶段的256点逆沃尔什-阿达玛变换精确逆变换FWHT,确保离线量化与在线推理间保持零误差往返保真度。我们证明,对于经本管线处理的任意权值向量$\mathbf{w} \in \mathbb{R}^{256}$,重构满足$\|\hat{\mathbf{w}} - \mathbf{w}\|_2 \leq ε_q$,其中$ε_q$仅由三值量化网格决定,且在等比特预算约束下严格优于任意均匀3位基线方法。实验表明,在NVIDIA RTX 5090(Blackwell架构)上,ITQ3\_S在实现与FP16基线相当的困惑度同时,因交错内存布局中优化的DP4A与Tensor Core调度,吞吐量超过4位替代方案的1.5倍。我们的研究将ITQ3\_S确立为面向消费级硬件高保真大语言模型部署的实用且具有数学基础解决方案。

0
下载
关闭预览

相关内容

Llama-3-SynE:实现有效且高效的大语言模型持续预训练
专知会员服务
36+阅读 · 2024年7月30日
超全总结:神经网络加速之量化模型 | 附带代码
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
13+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
VIP会员
最新内容
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
3+阅读 · 8月8日
《多域冲突比较支持模型》60页
专知会员服务
9+阅读 · 8月7日
面向2027年及未来的海军情报改革
专知会员服务
6+阅读 · 8月5日
相关VIP内容
Llama-3-SynE:实现有效且高效的大语言模型持续预训练
专知会员服务
36+阅读 · 2024年7月30日
相关基金
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
13+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员