We present a high-performance evaluation method for 4-center 2-particle integrals over Gaussian atomic orbitals with high angular momenta ($l\geq4$) and arbitrary contraction degrees on graphical processing units (GPUs) and other accelerators. The implementation uses the matrix form of McMurchie-Davidson recurrences. Evaluation of the 4-center integrals over four $l=6$ ($i$) Gaussian AOs in the double precision (FP64) on an NVIDIA V100 GPU outperforms the reference implementation of the Obara-Saika recurrences (${\tt Libint}$) running on a single Intel Xeon core by more than a factor of 1000, healthily exceeding the 73:1 ratio of the respective hardware peak FLOP rates while reaching almost 50\% of the V100 peak. The approach can be extended to support AOs with even higher angular momenta; for low angular momenta alternative approaches will be needed to achieve optimal performance. The implementation is part of an open-source ${\tt LibintX}$ library feely available at ${\tt github.com:ValeevGroup/LibintX}$.
翻译:我们提出了一种在图形处理单元(GPU)及其他加速器上高效计算高角动量($l\geq4$)和任意收缩度的高斯原子轨道四中心二粒子积分的方法。该实现采用了McMurchie-Davidson递推关系的矩阵形式。在NVIDIA V100 GPU上,对四个$l=6$($i$)高斯原子轨道进行双精度(FP64)四中心积分评估,其性能相比在单个Intel Xeon内核上运行的Obara-Saika递推关系参考实现(${\tt Libint}$)提高了超过1000倍,健康地超过了相应硬件峰值FLOP速率73:1的比率,同时达到了V100峰值性能的近50%。该方法可扩展以支持更高角动量的原子轨道;对于低角动量情况,则需要采用替代方法以实现最优性能。该实现是开源${\tt LibintX}$库的一部分,可在${\tt github.com/ValeevGroup/LibintX}$免费获取。