DRAM-based in-situ accelerators have shown their potential in addressing the memory wall challenge of the traditional von Neumann architecture. Such accelerators exploit charge sharing or logic circuits for simple logic operations at the DRAM subarray level. However, their throughput is limited due to low array utilization, as only a few row cells in a DRAM array participate in operations while most rows remain deactivated. Moreover, they require many cycles for more complex operations such as a multi-bit multiply-accumulate (MAC) operation, resulting in significant data access and movement and potentially worsening power efficiency. To overcome these limitations, this paper presents MAC-DO, an efficient and low-power DRAM-based in-situ accelerator. Compared to previous DRAM-based in-situ accelerators, a MAC-DO cell, consisting of two 1T1C DRAM cells (two transistors and two capacitors), innately supports a multi-bit MAC operation within a single cycle, ensuring good linearity and compatibility with existing 1T1C DRAM cells and array structures. This achievement is facilitated by a novel analog computation method utilizing charge steering. Additionally, MAC-DO enables concurrent individual MAC operations in each MAC-DO cell without idle cells, significantly improving throughput and energy efficiency. As a result, a MAC-DO array efficiently can accelerate matrix multiplications based on output stationary mapping, supporting the majority of computations performed in deep neural networks (DNNs). Furthermore, a MAC-DO array efficiently reuses three types of data (input, weight and output), minimizing data movement.
翻译:基于DRAM的原位加速器在解决传统冯·诺依曼架构的存储墙挑战方面展现出潜力。这类加速器利用电荷共享或逻辑电路在DRAM子阵列层级实现简单逻辑运算。然而,由于DRAM阵列中仅有少数行单元参与运算而大部分行保持未激活状态,其吞吐量受到低阵列利用率的限制。此外,对于多位乘累加(MAC)等更复杂运算,此类加速器需要大量时钟周期,导致显著的数据访问与搬运,并可能降低能效。为克服这些局限,本文提出MAC-DO——一种高效低功耗的DRAM原位加速器。与现有DRAM原位加速器相比,MAC-DO单元由两个1T1C DRAM单元(两个晶体管与两个电容器)构成,天然支持单周期内完成多位MAC运算,并确保与现有1T1C DRAM单元及阵列结构的良好线性度与兼容性。这一突破得益于采用电荷导向的新型模拟计算方法。此外,MAC-DO支持每个MAC-DO单元并发执行独立MAC运算,无闲置单元,从而显著提升吞吐量与能效。基于此,MAC-DO阵列能基于输出固定映射高效加速矩阵乘法,支撑深度神经网络(DNN)中的绝大多数计算。进一步地,MAC-DO阵列通过高效复用三种数据类型(输入、权重与输出),最大限度减少数据搬运。