Gaussian processes are a versatile probabilistic machine learning model whose effectiveness often depends on good hyperparameters, which are typically learned by maximising the marginal likelihood. In this work, we consider iterative methods, which use iterative linear system solvers to approximate marginal likelihood gradients up to a specified numerical precision, allowing a trade-off between compute time and accuracy of a solution. We introduce a three-level hierarchy of marginal likelihood optimisation for iterative Gaussian processes, and identify that the computational costs are dominated by solving sequential batches of large positive-definite systems of linear equations. We then propose to amortise computations by reusing solutions of linear system solvers as initialisations in the next step, providing a $\textit{warm start}$. Finally, we discuss the necessary conditions and quantify the consequences of warm starts and demonstrate their effectiveness on regression tasks, where warm starts achieve the same results as the conventional procedure while providing up to a $16 \times$ average speed-up among datasets.
翻译:高斯过程是一种多用途的概率机器学习模型,其有效性通常依赖于良好的超参数,这些超参数一般通过最大化边缘似然来学习。在本工作中,我们考虑迭代方法,该方法使用迭代线性系统求解器以指定的数值精度近似边缘似然梯度,从而在计算时间与解的精度之间进行权衡。我们为迭代高斯过程引入了一个三层级的边缘似然优化框架,并指出计算成本主要源于顺序求解多批大型正定线性方程组。我们进而提出通过在线性系统求解的下一步中复用前一步的解作为初始化来分摊计算成本,即提供一种$\textit{预热启动}$策略。最后,我们讨论了预热启动的必要条件,量化了其带来的影响,并在回归任务上验证了其有效性。实验表明,预热启动能达到与传统流程相同的结果,同时在多个数据集上实现了高达$16 \times$的平均加速比。