The rapid scaling of large language models (LLMs) exacerbates communication bottlenecks in AI data centers (AIDCs). To overcome this, optical circuit switches (OCS) are increasingly adopted for their superior bandwidth capacity and energy efficiency. However, their reconfiguration overhead precludes intra-iteration topology update, necessitating a priori engineering of a static topology to absorb time-varying LLM traffic. Existing methods engineer these topologies based on traffic matrices. However, this representation obscures the bursty concurrent bandwidth demands dictated by parallelization strategies and fails to account for the independent channels required for concurrent communication. To address this, we propose DELTA, an efficient logical topology optimization framework for AIDCs that leverages the computation-communication directed acyclic graph (DAG) to encode time-varying traffic patterns into a Mixed-Integer Linear Programming (MILP) model, while exploiting the temporal slack of non-critical tasks to save optical ports without penalizing iteration makespan. By pioneering a variable-length time interval formulation, DELTA significantly reduces the solution space compared to the fixed-time-step formulation. To scale to thousand-GPU clusters, we design a dual-track acceleration strategy that combines search space pruning (reducing complexity from quadratic to linear) with heuristic hot-starting. Evaluations on large-scale LLM workloads show that DELTA reduces communication time by up to 17.5% compared to state-of-the-art traffic-matrix-based baselines. Furthermore, the framework reduces optical port consumption by at least 20%; dynamically reallocating these surplus ports to bandwidth-bottlenecked workloads reduces their performance gap relative to ideal non-blocking electrical networks by up to 26.1%, ultimately enabling most workloads to achieve near-ideal performance.


翻译:大语言模型(LLM)的快速扩展加剧了AI数据中心(AIDC)中的通信瓶颈。为克服这一问题,光路交换机(OCS)因其优越的带宽容量和能效而被广泛采用。然而,其重构开销阻碍了迭代内的拓扑更新,需要预先设计静态拓扑以吸收时变的LLM流量。现有方法基于流量矩阵设计这些拓扑,但这种表示方式掩盖了由并行化策略决定的突发性并发带宽需求,且未能考虑并发通信所需的独立信道。为此,我们提出DELTA——一种高效的AIDC逻辑拓扑优化框架,该框架利用计算-通信有向无环图(DAG)将时变流量模式编码为混合整数线性规划(MILP)模型,同时利用非关键任务的时间松弛来节省光端口而不影响迭代总耗时。通过开创性地引入可变长度时间区间公式,DELTA相较于固定时间步长公式显著缩减了求解空间。为扩展到千GPU规模的集群,我们设计了一种双轨加速策略,将搜索空间剪枝(复杂度从二次降为线性)与启发式热启动相结合。在大规模LLM工作负载上的评估表明,与最先进的基于流量矩阵的基线方法相比,DELTA将通信时间减少了高达17.5%。此外,该框架至少减少了20%的光端口消耗;将这些多余端口动态重新分配给带宽瓶颈型工作负载,使其与理想非阻塞电网络的性能差距缩小了高达26.1%,最终使大多数工作负载达到近乎理想的性能。

0
下载
关闭预览

相关内容

大规模语言模型增强推荐系统:分类、趋势、应用与未来
专知会员服务
41+阅读 · 2024年12月22日
【学界】DeepMind论文:深度压缩感知,新框架提升GAN性能
GAN生成式对抗网络
14+阅读 · 2019年5月23日
CNN 模型压缩与加速算法综述
机器学习研究会
16+阅读 · 2017年8月25日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
相关主题
最新内容
《基于强化学习的自动化红队测试》
专知会员服务
3+阅读 · 7月23日
伊朗不对称防空战略的演进
专知会员服务
4+阅读 · 7月23日
对抗环境下超视距目标打击的情报支援
专知会员服务
10+阅读 · 7月22日
《无人机对海面作战影响评估》
专知会员服务
15+阅读 · 7月21日
印度精确打击与指挥架构的断层
专知会员服务
7+阅读 · 7月20日
相关VIP内容
大规模语言模型增强推荐系统:分类、趋势、应用与未来
专知会员服务
41+阅读 · 2024年12月22日
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员