Learning-in-memory (LIM) is a recently proposed paradigm to overcome fundamental memory bottlenecks in training machine learning systems. While compute-in-memory (CIM) approaches can address the so-called memory-wall (i.e. energy dissipated due to repeated memory read access) they are agnostic to the energy dissipated due to repeated memory writes at the precision required for training (the update-wall), and they don't account for the energy dissipated when transferring information between short-term and long-term memories (the consolidation-wall). The LIM paradigm proposes that these bottlenecks, too, can be overcome if the energy barrier of physical memories is adaptively modulated such that the dynamics of memory updates and consolidation match the Lyapunov dynamics of gradient-descent training of an AI model. In this paper, we derive new theoretical lower bounds on energy dissipation when training AI systems using different LIM approaches. The analysis presented here is model-agnostic and highlights the trade-off between energy efficiency and the speed of training. The resulting non-equilibrium energy-efficiency bounds have a similar flavor as that of Landauer's energy-dissipation bounds. We also extend these limits by taking into account the number of floating-point operations (FLOPs) used for training, the size of the AI model, and the precision of the training parameters. Our projections suggest that the energy-dissipation lower-bound to train a brain scale AI system (comprising of $10^{15}$ parameters) using LIM is $10^8 \sim 10^9$ Joules, which is on the same magnitude the Landauer's adiabatic lower-bound and $6$ to $7$ orders of magnitude lower than the projections obtained using state-of-the-art AI accelerator hardware lower-bounds.
翻译:内存学习(LIM)是近期提出的一种旨在克服机器学习系统训练中固有内存瓶颈的新范式。尽管内存计算(CIM)方法能够解决所谓的"内存墙"问题(即因重复内存读取访问而耗散的能量),但它们忽视了训练所需精度下因重复内存写入而耗散的能量("更新墙"),也未考虑在短期记忆与长期记忆之间传输信息所耗散的能量("巩固墙")。LIM范式提出,若能自适应地调制物理存储器的能量势垒,使得存储器更新与巩固的动态过程匹配人工智能模型梯度下降训练的Lyapunov动态,则这些瓶颈亦可被克服。本文推导了采用不同LIM方法训练人工智能系统时能量耗散的新理论下界。所提出的分析具有模型无关性,并揭示了能效与训练速度之间的权衡关系。所得非平衡态能效界限在形式上与Landauer能量耗散界限相似。我们进一步通过考虑训练所用的浮点运算次数(FLOPs)、人工智能模型的规模以及训练参数的精度,扩展了这些极限。我们的预测表明,采用LIM训练大脑规模人工智能系统(包含$10^{15}$个参数)的能量耗散下界为$10^8 \sim 10^9$焦耳,该数值与Landauer绝热下界处于同一数量级,且比基于现有最先进人工智能加速器硬件下界获得的预测值低$6$至$7$个数量级。