Communication is pivotal in LLM training, and a thorough analysis of the communication efficiency of AI data center (AIDC) network is essential for guiding the design of these capital-intensive clusters. However, conventional metrics are inadequate for such analysis, as they do not directly link network activity to computational progress and lack granularity to diagnose the impact of different network design patterns. To address this, we introduce a metric framework, the Switching Efficiency Framework, whose core metric - Switching Efficiency ($η$) - quantifies computationally effective data throughput per unit switching capacity. We further decompose $η$ into three factors - Data, Routing Efficiency, and Port Utilization to facilitate analysis of distinct communication bottlenecks. Using this metric framework, we demonstrate how the symmetric, distributed switching of 3D-Torus and the centralized, hierarchical switching of Rail-Optimized architecture align with sparse or imbalanced LLM training traffic, and show that All-to-All traffic from Mixture-of-Experts models severely degrades their port utilization and routing efficiency. Our analysis also demonstrates how key design choices - such as adjusting switching resource allocation, expanding server size, adopting in-network computing, and multi-plane design - positively influence distinct facets of communication efficiency. Ultimately, the Switching Efficiency Framework provides an analytical tool for analyzing efficiency bottlenecks, thereby informing the design of future-generation AIDC networks.


翻译:通信在大语言模型训练中至关重要,而对AI数据中心网络通信效率的深入分析是指导这些资本密集型集群设计的关键。然而,传统指标不足以支撑此类分析,因为它们无法直接将网络活动与计算进展相关联,且缺乏细粒度来诊断不同网络设计模式的影响。为此,我们提出了一种指标框架——切换效率框架,其核心指标——切换效率(η)——量化了单位交换容量下的计算有效数据吞吐量。我们进一步将η分解为三个因子——数据、路由效率和端口利用率,以便分析不同的通信瓶颈。利用该指标框架,我们展示了3D-Torus的对称分布式交换和Rail-Optimized架构的集中式分层交换如何适应稀疏或不均衡的大语言模型训练流量,并表明混合专家模型产生的全到全流量严重降低了它们的端口利用率和路由效率。我们的分析还展示了关键设计选择——如调整交换资源分配、扩展服务器规模、采用网内计算和多平面设计——如何积极影响通信效率的不同方面。最终,切换效率框架为分析效率瓶颈提供了一种分析工具,从而为未来一代AI数据中心网络的设计提供参考。

0
下载
关闭预览

相关内容

面向AI大模型的智算中心网络演进白皮书,30页pdf
专知会员服务
85+阅读 · 2023年5月15日
重磅!AI框架发展白皮书(2022年),44页pdf
专知
28+阅读 · 2022年2月27日
谷歌EfficientNet缩放模型,PyTorch实现登热榜
机器学习算法与Python学习
11+阅读 · 2019年6月4日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
VIP会员
相关主题
最新内容
乌克兰纵深打击如何重塑俄罗斯的战略选择
专知会员服务
1+阅读 · 今天12:25
俄乌战争中关于中程打击无人机部署的经验启示
专知会员服务
0+阅读 · 今天12:08
《基于强化学习的自动化红队测试》
专知会员服务
4+阅读 · 7月23日
伊朗不对称防空战略的演进
专知会员服务
4+阅读 · 7月23日
对抗环境下超视距目标打击的情报支援
专知会员服务
10+阅读 · 7月22日
相关VIP内容
面向AI大模型的智算中心网络演进白皮书,30页pdf
专知会员服务
85+阅读 · 2023年5月15日
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员