Additive Gaussian Processes (GPs) are popular approaches for nonparametric feature selection. The common training method for these models is Bayesian Back-fitting. However, the convergence rate of Back-fitting in training additive GPs is still an open problem. By utilizing a technique called Kernel Packets (KP), we prove that the convergence rate of Back-fitting is no faster than $(1-\mathcal{O}(\frac{1}{n}))^t$, where $n$ and $t$ denote the data size and the iteration number, respectively. Consequently, Back-fitting requires a minimum of $\mathcal{O}(n\log n)$ iterations to achieve convergence. Based on KPs, we further propose an algorithm called Kernel Multigrid (KMG). This algorithm enhances Back-fitting by incorporating a sparse Gaussian Process Regression (GPR) to process the residuals subsequent to each Back-fitting iteration. It is applicable to additive GPs with both structured and scattered data. Theoretically, we prove that KMG reduces the required iterations to $\mathcal{O}(\log n)$ while preserving the time and space complexities at $\mathcal{O}(n\log n)$ and $\mathcal{O}(n)$ per iteration, respectively. Numerically, by employing a sparse GPR with merely 10 inducing points, KMG can produce accurate approximations of high-dimensional targets within 5 iterations.
翻译:加性高斯过程是高斯基模型在非参数特征选择中的常用方法。此类模型的常见训练方法是贝叶斯反向拟合。然而,训练加性高斯过程时反向拟合方法的收敛速率仍是一个开放问题。通过利用称为核包的技术,我们证明了反向拟合的收敛速率不会快于$(1-\mathcal{O}(\frac{1}{n}))^t$,其中$n$和$t$分别表示数据规模和迭代次数。因此,反向拟合需要至少$\mathcal{O}(n\log n)$次迭代才能达到收敛。基于KP理论,我们进一步提出了一种名为核多网格的算法。该算法通过在每次反向拟合迭代后引入稀疏高斯过程回归处理残差,从而增强反向拟合性能。该算法适用于结构化和散乱数据下的加性高斯过程。理论上,我们证明KMG可将所需迭代次数减少至$\mathcal{O}(\log n)$,同时保持每次迭代的时间和空间复杂度分别为$\mathcal{O}(n\log n)$和$\mathcal{O}(n)$。数值实验表明,采用仅含10个诱导点的稀疏GPR时,KMG能在5次迭代内获得高维目标的精确近似。