Processing-in-memory (PIM), as a novel computing paradigm, provides significant performance benefits from the aspect of effective data movement reduction. SRAM-based PIM has been demonstrated as one of the most promising candidates due to its endurance and compatibility. However, the integration density of SRAM-based PIM is much lower than other non-volatile memory-based ones, due to its inherent 6T structure for storing a single bit. Within comparable area constraints, SRAM-based PIM exhibits notably lower capacity. Thus, aiming to unleash its capacity potential, we propose DDC-PIM, an efficient algorithm/architecture co-design methodology that effectively doubles the equivalent data capacity. At the algorithmic level, we propose a filter-wise complementary correlation (FCC) algorithm to obtain a bitwise complementary pair. At the architecture level, we exploit the intrinsic cross-coupled structure of 6T SRAM to store the bitwise complementary pair in their complementary states ($Q/\overline{Q}$), thereby maximizing the data capacity of each SRAM cell. The dual-broadcast input structure and reconfigurable unit support both depthwise and pointwise convolution, adhering to the requirements of various neural networks. Evaluation results show that DDC-PIM yields about $2.84\times$ speedup on MobileNetV2 and $2.69\times$ on EfficientNet-B0 with negligible accuracy loss compared with PIM baseline implementation. Compared with state-of-the-art SRAM-based PIM macros, DDC-PIM achieves up to $8.41\times$ and $2.75\times$ improvement in weight density and area efficiency, respectively.
翻译:存内计算(PIM)作为一种新型计算范式,通过有效减少数据移动显著提升了性能。基于SRAM的PIM因其耐久性和兼容性已被证明是最有前景的方案之一。然而,由于SRAM基本存储单元固有的6T结构,其集成密度远低于其他基于非易失性存储器的PIM方案。在同等面积约束下,SRAM型PIM的等效数据容量明显偏低。为此,我们提出DDC-PIM——一种高效算法/架构协同设计方法,可有效实现数据容量倍增。在算法层面,我们提出滤波器级互补相关(FCC)算法以生成按位互补对;在架构层面,利用6T SRAM固有的交叉耦合结构,将互补对存储于其互补状态($Q/\overline{Q}$),从而最大化每个SRAM单元的存储容量。双广播输入结构与可重构单元支持深度卷积和逐点卷积,满足各类神经网络需求。评估结果表明,与基准PIM实现相比,DDC-PIM在MobileNetV2和EfficientNet-B0上分别实现约$2.84\times$和$2.69\times$的加速比,精度损失可忽略。与当前最先进的SRAM型PIM宏单元相比,DDC-PIM在权重密度和面积效率上分别提升达$8.41\times$和$2.75\times$。