The cryosphere plays a significant role in Earth's climate system. Therefore, an accurate simulation of sea ice is of great importance to improve climate projections. To enable higher resolution simulations, graphics processing units (GPUs) have become increasingly attractive as they offer higher floating point peak performance and better energy efficiency compared to CPUs. However, making use of the theoretical peak performance usually requires more care and effort in the implementation. In recent years, a number of frameworks have become available that promise to simplify general purpose GPU programming. In this work, we compare multiple such frameworks, including CUDA, SYCL, Kokkos and PyTorch, for the parallelization of \nextsim, a finite-element based dynamical core for sea ice. We evaluate the different approaches according to their usability and performance.
翻译:冰冻圈在地球气候系统中扮演着重要角色。因此,对海冰进行精确模拟对于改进气候预测具有重要意义。为了实现更高分辨率的模拟,图形处理器(GPU)因其相较于CPU具有更高的浮点峰值性能和更好的能效而日益受到青睐。然而,要充分发挥理论峰值性能,通常需要在实现过程中投入更多精力与细致考量。近年来,多种旨在简化通用GPU编程的框架相继问世。本研究比较了CUDA、SYCL、Kokkos与PyTorch等框架在基于有限元的海冰动力核心neXtSIM并行化中的应用效果,并从可用性与性能两个维度对上述方案进行了评估。