Model size and inference speed at deployment time, are major challenges in many deep learning applications. A promising strategy to overcome these challenges is quantization. However, a straightforward uniform quantization to very low precision can result in significant accuracy loss. Mixed-precision quantization, based on the idea that certain parts of the network can accommodate lower precision without compromising performance compared to other parts, offers a potential solution. In this work, we present High Granularity Quantization (HGQ), an innovative quantization-aware training method designed to fine-tune the per-weight and per-activation precision in an automatic way for ultra-low latency and low power neural networks which are to be deployed on FPGAs. We demonstrate that HGQ can outperform existing methods by a substantial margin, achieving resource reduction by up to a factor of 20 and latency improvement by a factor of 5 while preserving accuracy.
翻译:模型大小与部署时的推理速度是众多深度学习应用面临的主要挑战。量化是一种有前景的应对策略,然而,简单地将所有参数统一量化为极低精度可能导致显著的精度损失。混合精度量化基于网络某些部分可承受更低精度而不影响整体性能的思想,为此提供了潜在解决方案。本文提出高粒度量化(HGQ)方法,这是一种创新的量化感知训练技术,旨在以自动方式精细调整每个权值与每个激活值的精度,适用于部署在FPGA上的超低延迟与低功耗神经网络。我们证明,HGQ能够在保持精度的前提下,以显著优势超越现有方法,实现资源消耗降低高达20倍、延迟提升5倍的性能。