Provisioning to Runtime Optimization of a +100 MW AI Cluster

Ehsan K. Ardestani,Leonardo Piga,Jovan Stojkovic,Pavan Balaji,Mustafa Ozdal,Mikel Jimenez Fernandez,Mihaela Dimovska,Luka Tadic,Hao Shen,Devika Vishwanath,Richa Mishra,Melaku Mihret,Valentin Andrei,Mauricio Cespedes,Julien Prigent,James Monahan,Tyler Graf,Bin Li,Charles Marquez,Shobhit Kanaujia,Kaushik Veeraraghavan,Chunqiang Tang

The electric power supply for AI data centers is now the most significant bottleneck in the race toward Artificial General Intelligence, surpassing even the constraint of AI accelerator availability. To our knowledge, this paper is the first to describe the end-to-end power management process for a hyper-scale AI datacenter; from early power planning to accommodate next-generation accelerators 6--12 months before their general availability, to tuning power settings after large scale deployment, and finally to dynamic, runtime power management for evolving workloads. We present detailed power measurements for a 150 MW datacenter hosting a cluster of 83K GB200 GPUs. We share insights from building this state-of-the-art AI cluster. We hope this work encourages practitioners across the industry to share their own experiences as well.

翻译：人工智能数据中心的电力供应已成为实现通用人工智能（AGI）的最大瓶颈，甚至超过AI加速器可用性的约束。据我们所知，本文首次描述了超大规模AI数据中心的端到端供电管理流程：从为下一代加速器实现提前6-12个月的早期电力规划，到大规模部署后的功耗设置调优，直至针对动态演变的运行负载进行运行时电力管理。我们展示了容纳83K GB200 GPU集群的150 MW数据中心的详细功耗测量数据，并分享了构建该先进AI集群的实践经验。希望这项研究能激励业界从业者共同分享自己的实践经验。

相关内容

关注 7110

人工智能杂志AI(Artificial Intelligence)是目前公认的发表该领域最新研究成果的主要国际论坛。该期刊欢迎有关AI广泛方面的论文，这些论文构成了整个领域的进步，也欢迎介绍人工智能应用的论文，但重点应该放在新的和新颖的人工智能方法如何提高应用领域的性能，而不是介绍传统人工智能方法的另一个应用。关于应用的论文应该描述一个原则性的解决方案，强调其新颖性，并对正在开发的人工智能技术进行深入的评估。官网地址：http://dblp.uni-trier.de/db/journals/ai/

《更智能的边缘：边缘计算如何推动美国人工智能领导力与能源安全》报告

专知会员服务

18+阅读 · 3月7日

下一代军事行动：利用人工智能提升后勤效率、能源管理与气候战备

专知会员服务

19+阅读 · 1月23日

EdgeRunner AI：在本地设备关键军事任务中实现GPT-5级性能表现（附论文）

专知会员服务

29+阅读 · 2025年11月19日

《面向边缘智能应用的AI模型优化技术研究》139页

专知会员服务

43+阅读 · 2025年8月12日