SoCs are now designed with their own AI accelerator segment to accommodate the ever-increasing demand of Deep Learning (DL) applications. With powerful MAC engines for matrix multiplications, these accelerators show high computing performance. However, because of limited memory resources (i.e., bandwidth and capacity), they fail to achieve optimum system performance during large batch training and inference. In this work, we propose a memory system with high on-chip capacity and bandwidth to shift the gear of AI accelerators from memory-bound to achieving system-level peak performance. We develop the memory system with DTCO-enabled customized SOT-MRAM as large on-chip memory through STCO and detailed characterization of the DL workloads. %We evaluate our workload-aware memory system on the CV and NLP benchmarks and observe significant PPA improvement compared to an SRAM-based in both inference and training modes. Our workload-aware memory system achieves 8X energy and 9X latency improvement on Computer Vision (CV) benchmarks in training and 8X energy and 4.5X latency improvement on Natural Language Processing (NLP) benchmarks in training while consuming only around 50% of SRAM area at iso-capacity.
翻译:SoC现配备专用AI加速器模块以应对深度学习(DL)应用日益增长的需求。这些加速器凭借用于矩阵乘法的强大MAC引擎展现出高性能计算能力。然而,由于内存资源(即带宽和容量)的限制,在大批量训练和推理过程中难以达到最优系统性能。本研究提出一种具备高片上容量和带宽的内存系统,使AI加速器从内存受限状态转变为实现系统级峰值性能。我们通过STCO方法开发采用DTCO定制的SOT-MRAM作为大容量片上内存的内存系统,并结合深度学习工作负载的详细特征分析。我们基于计算机视觉(CV)和自然语言处理(NLP)基准测试对工作负载感知内存系统进行评估,发现相比基于SRAM的方案,在推理和训练模式下均获得显著PPA改进。该工作负载感知内存系统在CV基准测试中实现训练阶段8倍能效提升和9倍延迟优化,在NLP基准测试中实现8倍能效提升和4.5倍延迟优化,同时等容量下占用面积仅为SRAM的约50%。