Recognizing the explosive increase in the use of DNN-based applications, several industrial companies developed a custom ASIC (e.g., Google TPU, IBM RaPiD, Intel NNP-I/NNP-T) and constructed a hyperscale cloud infrastructure with it. The ASIC performs operations of the inference or training process of DNN models which are requested by users. Since the DNN models have different data formats and types of operations, the ASIC needs to support diverse data formats and generality for the operations. However, the conventional ASICs do not fulfill these requirements. To overcome the limitations of it, we propose a flexible DNN accelerator called All-rounder. The accelerator is designed with an area-efficient multiplier supporting multiple precisions of integer and floating point datatypes. In addition, it constitutes a flexibly fusible and fissionable MAC array to support various types of DNN operations efficiently. We implemented the register transfer level (RTL) design using Verilog and synthesized it in 28nm CMOS technology. To examine practical effectiveness of our proposed designs, we designed two multiply units and three state-of-the-art DNN accelerators. We compare our multiplier with the multiply units and perform architectural evaluation on performance and energy efficiency with eight real-world DNN models. Furthermore, we compare benefits of the All-rounder accelerator to a high-end GPU card, i.e., NVIDIA GeForce RTX30390. The proposed All-rounder accelerator universally has speedup and high energy efficiency in various DNN benchmarks than the baselines.
翻译:摘要:认识到基于DNN的应用呈爆炸式增长,多家工业公司开发了定制ASIC(例如谷歌TPU、IBM RaPiD、英特尔NNP-I/NNP-T),并以此构建了超大规模云基础设施。该ASIC执行用户请求的DNN模型推理或训练过程的操作。由于DNN模型具有不同的数据格式和操作类型,ASIC需要支持多样化的数据格式和操作通用性。然而,传统ASIC无法满足这些要求。为克服其局限性,我们提出一种名为“全能手”的灵活DNN加速器。该加速器采用支持整数和浮点数据类型多精度的高面积效率乘法器设计。此外,它构建了灵活可融合与可分裂的MAC阵列,以高效支持各类DNN操作。我们使用Verilog实现了寄存器传输级(RTL)设计,并在28nm CMOS工艺下完成了综合。为验证所提设计的实际效果,我们设计了两种乘法单元和三种最先进的DNN加速器。我们将所提乘法器与这些乘法单元进行对比,并使用八种真实DNN模型进行架构层面的性能和能效评估。此外,我们将全能手加速器的优势与高端GPU显卡(即NVIDIA GeForce RTX3090)进行比较。所提出的全能手加速器在各种DNN基准测试中相比基线普遍具有更高的加速比和能效。