As the demand for deep learning grows, cost reduction through quantization has become essential for both training and inference. In 2022, the Open Compute Project (OCP) consortium standardized narrow precision formats for deep learning, called the microscaling (MX) format. The MX format is a hardware-friendly dynamic quantization scheme that effectively reduces the data size by sharing an 8-bit exponent across multiple operands. The MX format can be categorized into two types with their own strengths: (i) MXINT which focuses on a high precision consisting only of mantissa bits and (ii) MXFP which focuses on a wider dynamic range by allowing local exponent bits. In this work, we present a versatile MXFP format, called MX-SAFE (MXSF in short), that adaptively uses two modes, i.e., a wider mantissa mode (FP8 E2M5) and a subnormal FP mode (FP5 E3M2), to support both training and direct-cast inference. Furthermore, we propose a tile-based block design to increase hardware efficiency by reducing the burden of re-quantization process during the training with the MXSF format. Owing to the use of the proposed MXSF format, 0.05%/11.1% and 3.55%/3.57% improvements in accuracy, on average, for inference/full-training compared to MXFP8 E2M5 and MXFP8 E4M3 are observed, respectively. Moreover, we present a training-inference accelerator that supports the MXSF format and it achieves similar accuracy to the BF16 baseline while using 24.9% less total energy consumption.


翻译:随着深度学习需求的增长,通过量化降低计算成本已成为训练和推理过程中不可或缺的技术。2022年,开放计算项目(OCP)联盟为深度学习标准化了窄精度格式,即微缩放(MX)格式。该格式是一种硬件友好的动态量化方案,通过跨多个操作数共享8位指数有效减少数据规模。MX格式可分为两类,各具优势:(i)MXINT仅包含尾数位以实现高精度,(ii)MXFP通过允许本地指数位实现更广泛的动态范围。本文提出一种多功能MXFP格式——MX-SAFE(简称MXSF),该格式自适应采用两种模式:宽尾数模式(FP8 E2M5)和次正规FP模式(FP5 E3M2),同时支持训练与直接转换推理。此外,我们提出基于分块的块设计,通过减少MXSF格式训练过程中重量化过程的负担提升硬件效率。采用所提出的MXSF格式后,在推理/全训练任务中,相比MXFP8 E2M5和MXFP8 E4M3,平均准确率分别提升0.05%/11.1%和3.55%/3.57%。同时,我们提出支持MXSF格式的训练-推理加速器,其在总能耗降低24.9%的条件下实现了与BF16基线相当的准确率。

0
下载
关闭预览

相关内容

谷歌EfficientNet缩放模型,PyTorch实现登热榜
机器学习算法与Python学习
11+阅读 · 2019年6月4日
【学界】DeepMind论文:深度压缩感知,新框架提升GAN性能
GAN生成式对抗网络
14+阅读 · 2019年5月23日
超全总结:神经网络加速之量化模型 | 附带代码
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
9+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Arxiv
0+阅读 · 5月12日
VIP会员
最新内容
美海军陆战队将三型无人机整合入统一战场网络
专知会员服务
1+阅读 · 18分钟前
《无人机蜂群:释放人类-蜂群编队的潜能》
专知会员服务
1+阅读 · 45分钟前
《战略战术化:一项综合性述评》
专知会员服务
0+阅读 · 49分钟前
美陆军-工业界协同推进反无人机系统技术发展
专知会员服务
1+阅读 · 今天8:46
《跨域指挥背景下的领导力发展》最新报告
专知会员服务
1+阅读 · 今天8:40
俄乌无人机战争的六大启示
专知会员服务
10+阅读 · 8月3日
《无人机空中监控:通信实验洞察》
专知会员服务
8+阅读 · 8月3日
从采集到决策:美军视角下的战术情报范式重构
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
9+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员