Communication is pivotal in LLM training, and a thorough analysis of the communication efficiency of AI data center (AIDC) network is essential for guiding the design of these capital-intensive clusters. However, conventional metrics are inadequate for such analysis, as they do not directly link network activity to computational progress and lack granularity to diagnose the impact of different network design patterns. To address this, we introduce a metric framework, the Switching Efficiency Framework, whose core metric - Switching Efficiency ($η$) - quantifies computationally effective data throughput per unit switching capacity. We further decompose $η$ into three factors - Data, Routing Efficiency, and Port Utilization to facilitate analysis of distinct communication bottlenecks. Using this metric framework, we demonstrate how the symmetric, distributed switching of 3D-Torus and the centralized, hierarchical switching of Rail-Optimized architecture align with sparse or imbalanced LLM training traffic, and show that All-to-All traffic from Mixture-of-Experts models severely degrades their port utilization and routing efficiency. Our analysis also demonstrates how key design choices - such as adjusting switching resource allocation, expanding server size, adopting in-network computing, and multi-plane design - positively influence distinct facets of communication efficiency. Ultimately, the Switching Efficiency Framework provides an analytical tool for analyzing efficiency bottlenecks, thereby informing the design of future-generation AIDC networks.
翻译:通信在大语言模型训练中至关重要,而对AI数据中心网络通信效率的深入分析是指导这些资本密集型集群设计的关键。然而,传统指标不足以支撑此类分析,因为它们无法直接将网络活动与计算进展相关联,且缺乏细粒度来诊断不同网络设计模式的影响。为此,我们提出了一种指标框架——切换效率框架,其核心指标——切换效率(η)——量化了单位交换容量下的计算有效数据吞吐量。我们进一步将η分解为三个因子——数据、路由效率和端口利用率,以便分析不同的通信瓶颈。利用该指标框架,我们展示了3D-Torus的对称分布式交换和Rail-Optimized架构的集中式分层交换如何适应稀疏或不均衡的大语言模型训练流量,并表明混合专家模型产生的全到全流量严重降低了它们的端口利用率和路由效率。我们的分析还展示了关键设计选择——如调整交换资源分配、扩展服务器规模、采用网内计算和多平面设计——如何积极影响通信效率的不同方面。最终,切换效率框架为分析效率瓶颈提供了一种分析工具,从而为未来一代AI数据中心网络的设计提供参考。