This paper introduces a novel optimization framework for deep neural network (DNN) hardware accelerators, enabling the rapid development of customized and automated design flows. More specifically, our approach aims to automate the selection and configuration of low-level optimization techniques, encompassing DNN and FPGA low-level optimizations. We introduce novel optimization and transformation tasks for building design-flow architectures, which are highly customizable and flexible, thereby enhancing the performance and efficiency of DNN accelerators. Our results demonstrate considerable reductions of up to 92\% in DSP usage and 89\% in LUT usage for two networks, while maintaining accuracy and eliminating the need for human effort or domain expertise. In comparison to state-of-the-art approaches, our design achieves higher accuracy and utilizes three times fewer DSP resources, underscoring the advantages of our proposed framework.
翻译:本文提出了一种针对深度神经网络(DNN)硬件加速器的新型优化框架,实现了快速开发定制化且自动化设计流程的能力。具体而言,我们的方法旨在自动化底层优化技术的选择与配置,涵盖DNN及FPGA的底层优化。我们引入了构建设计流程架构的新型优化与变换任务,这些任务具有高度可定制性与灵活性,从而提升DNN加速器的性能与效率。实验结果表明,对于两个网络,该方法在保持精度的同时,将DSP使用量降低高达92%,LUT使用量降低89%,且无需人工干预或领域专业知识。与现有最先进方法相比,我们的设计在实现更高精度的同时,DSP资源消耗仅为前者的三分之一,充分彰显了所提框架的优势。