Disaggregated memory is a promising approach that addresses the limitations of traditional memory architectures by enabling memory to be decoupled from compute nodes and shared across a data center. Cloud platforms have deployed such systems to improve overall system memory utilization, but performance can vary across workloads. High-performance computing (HPC) is crucial in scientific and engineering applications, where HPC machines also face the issue of underutilized memory. As a result, improving system memory utilization while understanding workload performance is essential for HPC operators. Therefore, learning the potential of a disaggregated memory system before deployment is a critical step. This paper proposes a methodology for exploring the design space of a disaggregated memory system. It incorporates key metrics that affect performance on disaggregated memory systems: memory capacity, local and remote memory access ratio, injection bandwidth, and bisection bandwidth, providing an intuitive approach to guide machine configurations based on technology trends and workload characteristics. We apply our methodology to analyze thirteen diverse workloads, including AI training, data analysis, genomics, protein, fusion, atomic nuclei, and traditional HPC bookends. Our methodology demonstrates the ability to comprehend the potential and pitfalls of a disaggregated memory system and provides motivation for machine configurations. Our results show that eleven of our thirteen applications can leverage injection bandwidth disaggregated memory without affecting performance, while one pays a rack bisection bandwidth penalty and two pay the system-wide bisection bandwidth penalty. In addition, we also show that intra-rack memory disaggregation would meet the application's memory requirement and provide enough remote memory bandwidth. }
翻译:分离式内存是一种有前景的方法,通过将内存与计算节点解耦并在数据中心内共享,克服了传统内存架构的限制。云平台已部署此类系统以提高整体系统内存利用率,但不同工作负载的性能可能存在差异。高性能计算在科学与工程应用中至关重要,然而高性能计算集群同样面临内存利用率不足的问题。因此,在理解工作负载性能的同时提升系统内存利用率对HPC运营者而言至关重要。为此,在部署分离式内存系统前评估其潜力是关键的步骤。本文提出了一种探索分离式内存系统设计空间的方法论。该方法整合了影响分离式内存系统性能的关键指标:内存容量、本地与远程内存访问比例、注入带宽以及对分带宽,提供了一种直观的途径,可根据技术趋势与工作负载特性指导机器配置。我们应用该方法分析了13种多样化工作负载,包括AI训练、数据分析、基因组学、蛋白质、聚变、原子核及传统HPC基准案例。我们的方法展示了理解分离式内存系统潜力与缺陷的能力,并为机器配置提供了动机。结果表明,13个应用中,有11个可利用注入带宽型分离式内存而不影响性能,1个受限于机架对分带宽开销,2个受限于系统级对分带宽开销。此外,我们还指出机架内内存分离可满足应用的内存需求并提供足够的远程内存带宽。