Modern edge AI workloads demand maximum energy efficiency, motivating the pursuit of analog Compute-in-Memory (CIM) architectures. Simultaneously, the popularity of Large-Language-Models (LLMs) drives the adoption of low-bit floating-point formats which prioritize dynamic range. However, the conventional direct-accumulation CIM accommodates floating-points by normalizing them to a shared widened fixed-point scale. Consequently, hardware resolution is dictated by the input's dynamic range rather than its precision, and energy consumption is dominated by the ADC. We address this limitation by introducing local normalization for each input, weight, and multiply-accumulate (MAC) output via a Gain-Ranging MAC (GR-MAC). Normalization overhead is handled by low-power digital logic, enabling the computationally expensive MAC operation to remain in the energy-efficient low-precision analog regime. Energy modelling shows that the addition of a gain-ranging Stage to the MAC enables a 4-bit increase in input dynamic range without increased energy consumption at a 35 dB SQNR standard. Additionally, the ADC resolution requirement becomes invariant to input distribution assumptions, allowing construction of an upper bound with a 1.5-bit reduction compared to the conventional lower bound. These results establish a pathway towards unlocking favourable energy scaling trends of analog CIM for modern AI workloads.
翻译:现代边缘AI工作负载要求最大能效,这推动了模拟存内计算(CIM)架构的发展。与此同时,大语言模型(LLMs)的普及推动了优先考虑动态范围的低比特浮点格式的采用。然而,传统的直接累加型CIM通过将浮点数归一化到共享的加宽定点尺度来兼容浮点运算。因此,硬件分辨率由输入的动态范围而非其精度决定,且能耗主要由模数转换器(ADC)主导。我们通过引入针对每个输入、权重和乘积累加(MAC)输出的局部归一化(通过增益可调MAC(GR-MAC)实现)来解决这一局限。归一化开销由低功耗数字逻辑处理,使得计算密集型的MAC操作得以保持在高效低精度模拟域中。能量建模表明,在MAC中增加增益调节级可在35 dB SQNR标准下实现输入动态范围提升4比特且不增加能耗。此外,ADC分辨率要求不再受输入分布假设影响,相较于传统下界,可构建具有1.5比特缩减的上界。这些结果为解锁模拟CIM面向现代AI工作负载的有利能效缩放趋势开辟了路径。