Modern exascale GPU- and APU-based systems provide multiple power and energy sensors, but differences in scope, update rate, timing, and filtering complicate the attribution of short-lived accelerator activity. This paper presents a methodology to characterize and correct these effects on Cray EX systems with AMD Instinct MI250X GPUs (Frontier) and MI300A APUs (Portage). Using controlled square-wave workloads, we quantify update intervals, delay, aliasing, and variability across up to 512 GPUs and 480 APUs with on-chip (rocm-smi/amd-smi) and off-chip Cray Power Management sensors. We reconstruct power from cumulative energy counters to achieve faster response times, validate it against on-chip, off-chip, and node-level sensors, and integrate the resulting streams into a Score-P/PAPI-based tool for time-aligned, phase-level attribution. Applied to rocHPL, rocHPL-MxP, and HPG-MxP, the method separates energy savings due to reduced runtime from changes in power. Mixed precision reduces node energy on Frontier by 79% for rocHPL-MxP and 31% for HPG-MxP, with similar trends on Portage. These results provide portable guidance for sensor validation and power-aware optimization on current and future exascale systems.
翻译:现代基于GPU和APU的百亿亿次计算系统提供了多种功耗与能量传感器,但其在监测范围、更新率、时序和滤波方面的差异,使得对短时加速器活动的归因变得复杂。本文提出了一种在搭载AMD Instinct MI250X GPU(Frontier)和MI300A APU(Portage)的Cray EX系统上表征并校正这些影响的方法。通过受控方波负载,我们量化了多达512个GPU和480个APU在片上(roc-smi/amd-smi)与片外Cray电源管理传感器下的更新间隔、延迟、混叠及变异性。我们利用累积能量计数器重构功率以实现更快的响应时间,将其与片上、片外及节点级传感器进行验证,并将生成的流集成到基于Score-P/PAPI的工具中,用于时间对齐的相位级归因。将该方法应用于rocHPL、rocHPL-MxP和HPG-MxP,可分离因运行时间缩短导致的能量节省与功率变化。混合精度分别使Frontier上rocHPL-MxP和HPG-MxP的节点能量降低79%和31%,而在Portage上观察到了类似趋势。这些结果可为当前及未来百亿亿次系统上的传感器验证和功耗感知优化提供可移植的指导。