In-network computing techniques, exemplified by NVLink Sharp (NVLS), offer a promising approach to addressing the communication bottlenecks in LLM inference by offloading collective operations, such as All-Reduce, to switches. However, the accelerator-centric architecture of NVLS suffers from two fundamental limitations: 1) it relies on GPU load instructions to trigger reduction operations, which means that the data reduced in the switch must be additionally transferred back to the initiating GPU rather than being broadcast directly, thereby introducing unnecessary communication overhead; 2) due to its architectural constraints, NVLS cannot offload operators that are not decomposable into memory-semantic instructions, such as the in-network quantization (INQ) proposed in this work. As a result, All-Reduce in NVLS must operate at FP16/BF16 precision, leading to substantial bandwidth waste.To address these limitations, we propose SCIN, the first switch-centric in-network architecture for shared-memory networks of AI accelerators, enabling both low-latency and high-bandwidth All-Reduce. Specifically, we introduce an in-switch accelerator (ISA) capable of initiating memory-semantic operations for in-network processing, together with a co-designed communication fabric that incurs negligible protocol overhead. By eliminating redundant data movement, SCIN delivers lower All-Reduce latency than NVLS. Moreover, by integrating a quantization module into the ISA, SCIN enables INQ for All-Reduce, reducing its precision to 8 bits and nearly doubling bandwidth with negligible accuracy loss. We also present a prototype of SCIN on a multi-FPGA system to demonstrate its feasibility and effectiveness. Experimental results show that our design accelerates All-Reduce by up to 8.7x for small messages and 3.8x for large messages, leading up to 1.74x faster TTFT and 1.34x faster TPOT on LLaMA-2 models.


翻译:网络计算技术(如NVLink Sharp,简称NVLS)通过将All-Reduce等集合通信操作卸载至交换机,为缓解大语言模型推理中的通信瓶颈提供了有效途径。然而,NVLS的加速器中心架构存在两项根本性缺陷:1)其依赖GPU加载指令触发归约操作,导致交换机的归约结果必须回传至发起GPU而非直接广播,由此引入不必要的通信开销;2)受限于架构约束,NVLS无法卸载无法分解为内存语义指令的算子(如本文提出的网络内量化技术)。因此,NVLS的All-Reduce必须保持FP16/BF16精度,造成显著的带宽浪费。为解决上述问题,本文提出SCIN——面向AI加速器共享内存网络的以交换机为中心的创新架构,实现低延迟、高带宽的All-Reduce。具体而言,我们设计了支持发起内存语义操作进行网络内处理的交换机内加速器,并协同设计了协议开销可忽略的通信网络。通过消除冗余数据移动,SCIN的All-Reduce延迟低于NVLS。此外,在ISA中集成量化模块后,SCIN使All-Reduce支持网络内量化技术,将精度降至8位并在精度损失可忽略的情况下实现近乎双倍的带宽提升。我们在多FPGA系统上构建SCIN原型以验证其可行性与有效性。实验结果表明,本设计对短消息的All-Reduce加速比最高达8.7倍,长消息达3.8倍,在LLaMA-2模型上TTFT(首个令牌生成时间)提升1.74倍,TPOT(每个输出令牌生成时间)提升1.34倍。

0
下载
关闭预览

相关内容

TransMLA:多头潜在注意力(MLA)即为所需
专知会员服务
23+阅读 · 2025年2月13日
LLMCad:快速可扩展的设备上大型语言模型推理
专知会员服务
35+阅读 · 2023年9月11日
神经网络加速器架构概述
专知会员服务
37+阅读 · 2022年4月23日
NetworkMiner - 网络取证分析工具
黑白之道
16+阅读 · 2018年6月29日
网络表示学习领域(NRL/NE)必读论文汇总
AI科技评论
16+阅读 · 2018年2月18日
从LeNet到SENet——卷积神经网络回顾
AI科技评论
13+阅读 · 2018年2月15日
ResNet, AlexNet, VGG, Inception:各种卷积网络架构的理解
全球人工智能
20+阅读 · 2017年12月17日
国家自然科学基金
1+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
VIP会员
相关主题
最新内容
致命七类无人机:无人机时代的演进型合成兵种
《异构无人水面艇集群作战自主制导算法》130页
相关基金
国家自然科学基金
1+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
Top
微信扫码咨询专知VIP会员