Processing-In-Memory (PIM) accelerators have the potential to efficiently run Deep Neural Network (DNN) inference by reducing costly data movement and by using resistive RAM (ReRAM) for efficient analog compute. Unfortunately, overall PIM accelerator efficiency is limited by energy-intensive analog-to-digital converters (ADCs). Furthermore, existing accelerators that reduce ADC cost do so by changing DNN weights or by using low-resolution ADCs that reduce output fidelity. These strategies harm DNN accuracy and/or require costly DNN retraining to compensate. To address these issues, we propose the RAELLA architecture. RAELLA adapts the architecture to each DNN; it lowers the resolution of computed analog values by encoding weights to produce near-zero analog values, adaptively slicing weights for each DNN layer, and dynamically slicing inputs through speculation and recovery. Low-resolution analog values allow RAELLA to both use efficient low-resolution ADCs and maintain accuracy without retraining, all while computing with fewer ADC converts. Compared to other low-accuracy-loss PIM accelerators, RAELLA increases energy efficiency by up to 4.9$\times$ and throughput by up to 3.3$\times$. Compared to PIM accelerators that cause accuracy loss and retrain DNNs to recover, RAELLA achieves similar efficiency and throughput without expensive DNN retraining.
翻译:处理进存储器(PIM)加速器通过减少昂贵的数据移动并利用电阻式RAM(ReRAM)进行高效模拟计算,有望高效运行深度神经网络(DNN)推理。然而,PIM加速器的整体效率受限于高能耗的模数转换器(ADC)。此外,现有降低ADC成本的加速器通常通过改变DNN权重或使用降低输出保真度的低分辨率ADC来实现。这些策略会损害DNN精度和/或需要昂贵的DNN重新训练来补偿。为解决这些问题,我们提出了RAELLA架构。RAELLA针对每个DNN调整架构:通过编码权重以产生接近零的模拟值来降低计算模拟值的分辨率;为每个DNN层自适应分割权重;并通过推测与恢复动态分割输入。低分辨率模拟值使RAELLA能够使用高效的低分辨率ADC并保持精度而无需重新训练,同时减少ADC转换次数。与其他低精度损失的PIM加速器相比,RAELLA的能效提升最高达4.9倍,吞吐量提升最高达3.3倍。与导致精度损失并通过重新训练DNN恢复的PIM加速器相比,RAELLA在无需昂贵的DNN重新训练的情况下实现了相当的能效与吞吐量。