Wave propagation based on the spectral element method (SEM) is a representative HPC workload, but existing SEM implementations are not well matched to emerging ARM multicore CPUs with Scalable Matrix Extension (SME). We present an SME-enabled optimization of \textsc{SPECFEM3D} on the emerging LX2 processor that combines an SME-aware batched small-matrix kernel for SEM tensor-product operators, a memory-aware hybrid MPI+OpenMP execution scheme for limited-HBM systems, and a dispersion-based iso-accuracy study of the $(h,p)$ tradeoff. At fixed polynomial order, the optimized implementation improves full-application performance by 4--6$\times$ over the original code and delivers clear gains over optimized non-SME CPU baselines. Beyond these implementation-level gains, our results suggest that SME shifts the performance-favorable operating point toward higher polynomial orders along the dispersion-based iso-accuracy frontier, further reducing time-to-solution and working-set size. These results indicate that SME affects not only kernel efficiency, but also the practical discretization tradeoff for SEM on modern ARM multicore platforms.
翻译:基于谱元法(SEM)的波传播计算是高性能计算(HPC)领域的典型负载,但现有SEM实现方案难以适配配备可扩展矩阵扩展(SME)的新兴ARM多核CPU。本文针对新兴LX2处理器提出了一种启用SME的优化方案,具体包括:面向SEM张量积算子的SME感知批处理小矩阵核、面向有限高带宽存储(HBM)系统的内存感知混合MPI+OpenMP执行框架,以及基于色散等精度分析的(h,p)权衡研究。在固定多项式阶数下,优化实现使完整应用性能较原始代码提升4-6倍,并相较优化后的非SME CPU基线方案展现显著优势。除执行效率提升外,研究结果揭示:沿色散等精度前沿,SME将性能优势工作点向更高多项式阶数偏移,进一步缩短求解时间并降低工作集规模。这表明SME不仅影响核心计算效率,更改变了现代ARM多核平台上SEM离散方案的实用权衡策略。