This paper introduces Block Data Representations (BDR), a framework for exploring and evaluating a wide spectrum of narrow-precision formats for deep learning. It enables comparison of popular quantization standards, and through BDR, new formats based on shared microexponents (MX) are identified, which outperform other state-of-the-art quantization approaches, including narrow-precision floating-point and block floating-point. MX utilizes multiple levels of quantization scaling with ultra-fine scaling factors based on shared microexponents in the hardware. The effectiveness of MX is demonstrated on real-world models including large-scale generative pretraining and inferencing, and production-scale recommendation systems.
翻译:本文提出块数据表示(BDR)框架,用于探索和评估深度学习中的宽频谱窄精度格式。该框架支持主流量化标准的比较,并通过BDR识别出基于共享微指数(MX)的新型格式,这些格式在性能上超越了包括窄精度浮点数和块浮点数在内的其他先进量化方法。MX利用硬件中基于共享微指数的超细粒度缩放因子,实现多级量化缩放。在包含大规模生成式预训练、推理以及生产级推荐系统的实际模型上,MX的有效性得到了充分验证。