Learning curve extrapolation aims to predict model performance in later epochs of training, based on the performance in earlier epochs. In this work, we argue that, while the inherent uncertainty in the extrapolation of learning curves warrants a Bayesian approach, existing methods are (i) overly restrictive, and/or (ii) computationally expensive. We describe the first application of prior-data fitted neural networks (PFNs) in this context. A PFN is a transformer, pre-trained on data generated from a prior, to perform approximate Bayesian inference in a single forward pass. We propose LC-PFN, a PFN trained to extrapolate 10 million artificial right-censored learning curves generated from a parametric prior proposed in prior art using MCMC. We demonstrate that LC-PFN can approximate the posterior predictive distribution more accurately than MCMC, while being over 10 000 times faster. We also show that the same LC-PFN achieves competitive performance extrapolating a total of 20 000 real learning curves from four learning curve benchmarks (LCBench, NAS-Bench-201, Taskset, and PD1) that stem from training a wide range of model architectures (MLPs, CNNs, RNNs, and Transformers) on 53 different datasets with varying input modalities (tabular, image, text, and protein data). Finally, we investigate its potential in the context of model selection and find that a simple LC-PFN based predictive early stopping criterion obtains 2 - 6x speed-ups on 45 of these datasets, at virtually no overhead.
翻译:学习曲线外推旨在根据训练早期阶段的模型性能预测后期阶段的性能。本文论证了,尽管学习曲线外推中的固有不确定性需要采用贝叶斯方法,但现有方法存在(i)过度约束和/或(ii)计算成本过高的问题。我们首次描述了在先验数据拟合神经网络(PFN)在该场景中的应用。PFN是一种基于先验生成数据预训练的Transformer,可通过单次前向传播实现近似贝叶斯推断。我们提出LC-PFN——一种PFN,其训练目标是用MCMC方法外推由现有文献提出的参数先验生成的1000万条人工右删失学习曲线。实验表明,LC-PFN不仅能比MCMC更准确地逼近后验预测分布,同时速度提升超过10000倍。我们还证明,同一LC-PFN在来自四个学习曲线基准(LCBench、NAS-Bench-201、Taskset和PD1)的共计20000条真实学习曲线外推任务中具有竞争力竞争力——这些曲线源于在53个不同数据集(涵盖表格、图像、文本和蛋白质数据四种输入模态)上训练多种模型架构(MLP、CNN、RNN和Transformer)的过程。最后,我们探索了该方法在模型选择中的潜力,发现基于LC-PFN的简单预测性早停准则可在其中45个数据集上实现2-6倍的加速,且几乎不增加额外计算开销。