We present our experience in porting optimized CUDA implementations to oneAPI. We focus on the use case of numerical integration, particularly the CUDA implementations of PAGANI and $m$-Cubes. We faced several challenges that caused performance degradation in the oneAPI ports. These include differences in utilized registers per thread, compiler optimizations, and mappings of CUDA library calls to oneAPI equivalents. After addressing those challenges, we tested both the PAGANI and m-Cubes integrators on numerous integrands of various characteristics. To evaluate the quality of the ports, we collected performance metrics of the CUDA and oneAPI implementations on the Nvidia V100 GPU. We found that the oneAPI ports often achieve comparable performance to the CUDA versions, and that they are at most 10% slower.
翻译:我们介绍了将优化的CUDA实现移植到oneAPI的经验。我们重点关注数值积分这一应用场景,特别是PAGANI和$m$-Cubes的CUDA实现。我们在移植过程中遇到了若干挑战,导致oneAPI移植版的性能下降。这些挑战包括每个线程使用的寄存器数量差异、编译器优化差异以及CUDA库调用到oneAPI等效调用的映射差异。在解决这些挑战后,我们使用多种不同特征的被积函数对PAGANI和m-Cubes积分器进行了测试。为了评估移植质量,我们在Nvidia V100 GPU上收集了CUDA和oneAPI实现的性能指标。我们发现,oneAPI移植版通常能达到与CUDA版本相当的性能,且性能最多慢10%。