Analog In-Memory Computing (AIMC) is an emerging technology for fast and energy-efficient Deep Learning (DL) inference. However, a certain amount of digital post-processing is required to deal with circuit mismatches and non-idealities associated with the memory devices. Efficient near-memory digital logic is critical to retain the high area/energy efficiency and low latency of AIMC. Existing systems adopt Floating Point 16 (FP16) arithmetic with limited parallelization capability and high latency. To overcome these limitations, we propose a Near-Memory digital Processing Unit (NMPU) based on fixed-point arithmetic. It achieves competitive accuracy and higher computing throughput than previous approaches while minimizing the area overhead. Moreover, the NMPU supports standard DL activation steps, such as ReLU and Batch Normalization. We perform a physical implementation of the NMPU design in a 14 nm CMOS technology and provide detailed performance, power, and area assessments. We validate the efficacy of the NMPU by using data from an AIMC chip and demonstrate that a simulated AIMC system with the proposed NMPU outperforms existing FP16-based implementations, providing 139$\times$ speed-up, 7.8$\times$ smaller area, and a competitive power consumption. Additionally, our approach achieves an inference accuracy of 86.65 %/65.06 %, with an accuracy drop of just 0.12 %/0.4 % compared to the FP16 baseline when benchmarked with ResNet9/ResNet32 networks trained on the CIFAR10/CIFAR100 datasets, respectively.
翻译:模拟存内计算(AIMC)是一种用于实现快速且高能效深度学习推理的新兴技术。然而,为应对电路失配和存储器件非理想特性,需要一定量的数字后处理。高效的近存数字逻辑对于保持AIMC的高面积/能效比和低延迟至关重要。现有系统采用并行化能力有限且延迟较高的16位浮点运算。为克服这些局限,我们提出一种基于定点运算的近存数字处理单元(NMPU)。该单元在最小化面积开销的同时,实现了优于先前方法的竞争性精度和更高计算吞吐量。此外,NMPU支持ReLU和批归一化等标准深度学习激活步骤。我们基于14纳米CMOS工艺完成了NMPU设计的物理实现,并提供了详细的性能、功耗和面积评估。通过使用AIMC芯片的数据验证NMPU的有效性,结果表明:采用所提NMPU的模拟AIMC系统优于现有基于FP16的实现方案,实现了139倍加速、7.8倍面积缩减以及具有竞争力的功耗。此外,在采用CIFAR10/CIFAR100数据集训练的ResNet9/ResNet32网络基准测试中,本方法分别实现了86.65%/65.06%的推理精度,与FP16基线相比精度损失仅为0.12%/0.4%。