Recent advancements in quantization and mixed-precision techniques offer significant promise for improving the run-time and energy efficiency of neural networks. In this work, we further showed that neural networks, wherein individual parameters or activations can take on different precisions ranging between 1 and 4 bits, can achieve accuracies comparable to or exceeding the full-precision counterparts. However, the deployment of such networks poses numerous challenges, stemming from the necessity to manage and control the compute/communication/storage requirements associated with these extremely fine-grained mixed precisions for each piece of data. There is a lack of existing efficient hardware and system-level support tailored to these unique and challenging requirements. Our research introduces the first novel holistic hardware-software co-design approach for these networks, which enables a continuous feedback loop between hardware design, training, and inference to facilitate systematic design exploration. As a proof-of-concept, we illustrate this co-design approach by designing new, configurable CPU SIMD architectures tailored for these networks, tightly integrating the architecture with new system-aware training and inference techniques. We perform systematic design space exploration using this framework to analyze various tradeoffs. The design for mixed-precision networks that achieves optimized tradeoffs corresponds to an architecture that supports 1, 2, and 4-bit fixed-point operations with four configurable precision patterns, when coupled with system-aware training and inference optimization -- networks trained for this design achieve accuracies that closely match full-precision accuracies, while compressing and improving run-time efficiency of the neural networks drastically by 10-20x, compared to full-precision networks.
翻译:近年来,量化与混合精度技术的进步为提升神经网络的运行时效率和能效带来了显著前景。本研究进一步表明,当单个参数或激活值可采用1至4比特范围内的不同精度时,神经网络能够达到与全精度网络相当甚至更优的准确率。然而,此类网络的部署面临诸多挑战,根源在于需针对每份数据管理并控制这些极细粒度混合精度所带来的计算/通信/存储需求。目前缺乏专门满足这些独特且苛刻要求的现有高效硬件与系统级支持。本研究首次提出针对此类网络的整体性新型软硬件协同设计方法,该方法在硬件设计、训练与推理之间建立连续反馈回路,以促进系统性设计空间探索。作为概念验证,我们通过为这些网络设计新型可配置CPU SIMD架构来阐释该协同设计方法,并将该架构与新型系统感知训练及推理技术紧密集成。我们利用该框架进行系统性设计空间探索,分析各类权衡。实现优化权衡的混合精度网络设计,对应于一种支持1、2、4比特定点运算且具有四种可配置精度模式的架构,当结合系统感知训练与推理优化时——针对该设计训练的网络能够达到与全精度网络高度接近的准确率,同时将神经网络的压缩比和运行时效率提升至全精度网络的10-20倍。