We present a high-performance evaluation method for 4-center 2-particle integrals over Gaussian atomic orbitals with high angular momenta ($l\geq4$) and arbitrary contraction degrees on graphical processing units (GPUs) and other accelerators. The implementation uses the matrix form of McMurchie-Davidson recurrences. Evaluation of the 4-center integrals over four $l=6$ ($i$) Gaussian AOs in the double precision (FP64) on an NVIDIA V100 GPU outperforms the reference implementation of the Obara-Saika recurrences (${\tt Libint}$) running on a single Intel Xeon core by more than a factor of 1000, easily exceeding the 73:1 ratio of the respective hardware peak FLOP rates while reaching almost 50% of the V100 peak. The approach can be extended to support AOs with even higher angular momenta; for lower angular momenta ($l\leq3$) additional improvements will be reported elsewhere. The implementation is part of an open-source ${\tt LibintX}$ library feely available at ${\tt github.com:ValeevGroup/LibintX}$.
翻译:我们提出了一种在图形处理单元(GPU)及其他加速器上高效计算高斯原子轨道高角动量($l\geq4$)及任意收缩阶数的四中心二粒子积分的方法。该实现采用McMurchie-Davidson递推公式的矩阵形式。在NVIDIA V100 GPU上,以双精度(FP64)计算四个$l=6$($i$)高斯原子轨道的四中心积分时,其性能比在单个Intel Xeon核心上运行的Obara-Saika递推公式参考实现(${\tt Libint}$)高出超过1000倍,轻松超越了73:1的硬件峰值FLOP比率,同时达到V100峰值性能的近50%。该方法可扩展以支持更高角动量的原子轨道;对于较低角动量($l\leq3$)的额外改进将在另文中报告。该实现是开源${\tt LibintX}$库的一部分,可在${\tt github.com/ValeevGroup/LibintX}$免费获取。