GPU computing is expected to play an integral part in all modern Exascale supercomputers. It is also expected that higher order Godunov schemes will make up about a significant fraction of the application mix on such supercomputers. It is, therefore, very important to prepare the community of users of higher order schemes for hyperbolic PDEs for this emerging opportunity. We focus on three broad and high-impact areas where higher order Godunov schemes are used. The first area is computational fluid dynamics (CFD). The second is computational magnetohydrodynamics (MHD) which has an involution constraint that has to be mimetically preserved. The third is computational electrodynamics (CED) which has involution constraints and also extremely stiff source terms. Together, these three diverse uses of higher order Godunov methodology, cover many of the most important applications areas. In all three cases, we show that the optimal use of algorithms, techniques and tricks, along with the use of OpenACC, yields superlative speedups on GPUs! As a bonus, we find a most remarkable and desirable result: some higher order schemes, with their larger operations count per zone, show better speedup than lower order schemes on GPUs. In other words, the GPU is an optimal stratagem for overcoming the higher computational complexities of higher order schemes! Several avenues for future improvement have also been identified. A scalability study is presented for a real-world application using GPUs and comparable numbers of high-end multicore CPUs. It is found that GPUs offer a substantial performance benefit over comparable number of CPUs, especially when all the methods designed in this paper are used.
翻译:GPU计算预计将在所有现代百亿亿次超级计算机中发挥不可或缺的作用,同时高阶Godunov格式预计将占此类超级计算机应用组合的重要部分。因此,引导高阶格式用户群体为这一新兴机遇做好准备至关重要。我们聚焦于采用高阶Godunov格式的三个广泛且高影响力的领域:第一个领域是计算流体动力学(CFD);第二个领域是计算磁流体动力学(MHD),其具有需仿射保持的卷积约束;第三个领域是计算电动力学(CED),不仅包含卷积约束,还涉及极刚性源项。这三个高阶Godunov方法的不同应用领域覆盖了最重要的应用场景。在所有三个案例中,我们展示出算法、技术、技巧的优化运用,结合OpenACC的使用,可在GPU上实现卓越的加速效果!更为重要的是,我们获得了一个显著且令人欣喜的结果:部分高阶格式(尽管每网格单元计算量更大)在GPU上展现比低阶格式更优的加速比。换言之,GPU是克服高阶格式更高计算复杂度的最优策略。我们还确定了未来改进的多个方向。针对实际应用,我们展示了使用GPU与同等数量高端多核CPU的可扩展性研究。结果表明,采用本文设计的所有方法时,GPU相比同等数量CPU具有显著的性能优势。