Stratum：面向高效MoE推理的层级化单片三维堆叠DRAM系统-硬件协同设计 (Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE Serving)

As Large Language Models (LLMs) continue to evolve, Mixture of Experts (MoE) architecture has emerged as a prevailing design for achieving state-of-the-art performance across a wide range of tasks. MoE models use sparse gating to activate only a handful of expert sub-networks per input, achieving billion-parameter capacity with inference costs akin to much smaller models. However, such models often pose challenges for hardware deployment due to the massive data volume introduced by the MoE layers. To address the challenges of serving MoE models, we propose Stratum, a system-hardware co-design approach that combines the novel memory technology Monolithic 3D-Stackable DRAM (Mono3D DRAM), near-memory processing (NMP), and GPU acceleration. The logic and Mono3D DRAM dies are connected through hybrid bonding, whereas the Mono3D DRAM stack and GPU are interconnected via silicon interposer. Mono3D DRAM offers higher internal bandwidth than HBM thanks to the dense vertical interconnect pitch enabled by its monolithic structure, which supports implementations of higher-performance near-memory processing. Furthermore, we tackle the latency differences introduced by aggressive vertical scaling of Mono3D DRAM along the z-dimension by constructing internal memory tiers and assigning data across layers based on access likelihood, guided by topic-based expert usage prediction to boost NMP throughput. The Stratum system achieves up to 8.29x improvement in decoding throughput and 7.66x better energy efficiency across various benchmarks compared to GPU baselines.

翻译：随着大型语言模型（LLM）的持续演进，混合专家（MoE）架构已成为在广泛任务中实现最先进性能的主流设计。MoE模型通过稀疏门控机制，仅针对每个输入激活少量专家子网络，从而在保持接近较小模型推理成本的同时实现数十亿参数规模。然而，此类模型因MoE层引入的海量数据往往给硬件部署带来挑战。为应对MoE模型推理的挑战，本文提出Stratum——一种融合新型存储技术单片三维堆叠DRAM（Mono3D DRAM）、近内存处理（NMP）与GPU加速的系统-硬件协同设计方案。逻辑芯片与Mono3D DRAM芯片通过混合键合互连，而Mono3D DRAM堆栈与GPU则通过硅中介层连接。得益于其单片结构实现的高密度垂直互连间距，Mono3D DRAM具备比HBM更高的内部带宽，为高性能近内存处理提供了硬件基础。此外，针对Mono3D DRAM沿z轴激进垂直缩放带来的延迟差异，我们构建了内部存储层级，并基于主题驱动的专家使用预测机制，依据数据访问概率将其分配至不同存储层，从而提升NMP吞吐量。实验表明，相较于GPU基线系统，Stratum在多种基准测试中实现了最高8.29倍的解码吞吐量提升与7.66倍的能效优化。

相关内容

MoDELS

关注 45

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

【NeurIPS2021】用于文本图表示学习的 GNN 嵌套 Transformer 模型：GraphFormers

专知会员服务

46+阅读 · 2021年11月24日

Linux导论，Introduction to Linux，96页ppt

专知会员服务

82+阅读 · 2020年7月26日

FlowQA: Grasping Flow in History for Conversational Machine Comprehension

专知会员服务

34+阅读 · 2019年10月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日