Data centers are increasingly using more energy due to the rise in Artificial Intelligence (AI) workloads, which negatively impacts the environment and raises operational costs. Reducing operating expenses and carbon emissions while maintaining performance in data centers is a challenging problem. This work introduces a unique approach combining Game Theory (GT) and Deep Reinforcement Learning (DRL) for optimizing the distribution of AI inference workloads in geo-distributed data centers to reduce carbon emissions and cloud operating (energy + data transfer) costs. The proposed technique integrates the principles of non-cooperative Game Theory into a DRL framework, enabling data centers to make intelligent decisions regarding workload allocation while considering the heterogeneity of hardware resources, the dynamic nature of electricity prices, inter-data center data transfer costs, and carbon footprints. We conducted extensive experiments comparing our game-theoretic DRL (GT-DRL) approach with current DRL-based and other optimization techniques. The results demonstrate that our strategy outperforms the state-of-the-art in reducing carbon emissions and minimizing cloud operating costs without compromising computational performance. This work has significant implications for achieving sustainability and cost-efficiency in data centers handling AI inference workloads across diverse geographic locations.
翻译:随着人工智能(AI)工作负载的激增,数据中心能源消耗日益增长,对生态环境造成负面影响并推高运营成本。如何在不牺牲计算性能的前提下,实现数据中心运营成本与碳排放量的协同优化,已成为亟待解决的难题。本文提出一种融合博弈论(Game Theory, GT)与深度强化学习(Deep Reinforcement Learning, DRL)的创新方法,用于优化地理分布式数据中心中AI推理工作负载的调度策略,以实现碳排放与云端运营(能耗+数据传输)成本的双重降低。该技术通过将非合作博弈论原理融入DRL框架,使数据中心能综合考虑硬件资源异构性、电价动态波动、跨数据中心数据传输成本及碳足迹等多维因素,做出智能化的负载分配决策。我们开展了大量对比实验,将所提出的博弈论-深度强化学习(GT-DRL)方法与当前基于DRL及其他优化技术的方案进行系统比较。结果表明,本策略在降低碳排放和最小化云端运营成本方面均优于当前最优方法,且计算性能未受损失。该研究对实现跨地理区域处理AI推理工作负载数据中心的可持续性与成本效益具有重要指导意义。