Kernel machines have sustained continuous progress in the field of quantum chemistry. In particular, they have proven to be successful in the low-data regime of force field reconstruction. This is because many equivariances and invariances due to physical symmetries can be incorporated into the kernel function to compensate for much larger datasets. So far, the scalability of kernel machines has however been hindered by its quadratic memory and cubical runtime complexity in the number of training points. While it is known, that iterative Krylov subspace solvers can overcome these burdens, their convergence crucially relies on effective preconditioners, which are elusive in practice. Effective preconditioners need to partially pre-solve the learning problem in a computationally cheap and numerically robust manner. Here, we consider the broad class of Nystr\"om-type methods to construct preconditioners based on successively more sophisticated low-rank approximations of the original kernel matrix, each of which provides a different set of computational trade-offs. All considered methods aim to identify a representative subset of inducing (kernel) columns to approximate the dominant kernel spectrum.
翻译:核机器在量子化学领域持续取得进展。特别是在力场重构的低数据量场景中,此类方法已被证明卓有成效,这是因为物理对称性导致的等变性和不变性可被纳入核函数设计中,从而补偿更大规模数据集的缺失。然而,当前核机器的可扩展性仍受制于其关于训练样本数量的二次存储复杂度与三次时间复杂度。尽管已知迭代Krylov子空间求解器可克服这些障碍,但其收敛性严重依赖实际中难以获取的有效预处理器。有效的预处理器需要以计算经济且数值鲁棒的方式部分预解学习问题。本文考虑采用Nyström型方法的广泛分类,基于逐次更精细的原始核矩阵低秩近似来构建预处理器——每种近似均提供不同的计算权衡方案。所有被考察方法均致力于识别具有代表性的诱导(核)列子集,以近似主导核谱特征。