We introduce TabRepo, a new dataset of tabular model evaluations and predictions. TabRepo contains the predictions and metrics of 1206 models evaluated on 200 regression and classification datasets. We illustrate the benefit of our datasets in multiple ways. First, we show that it allows to perform analysis such as comparing Hyperparameter Optimization against current AutoML systems while also considering ensembling at no cost by using precomputed model predictions. Second, we show that our dataset can be readily leveraged to perform transfer-learning. In particular, we show that applying standard transfer-learning techniques allows to outperform current state-of-the-art tabular systems in accuracy, runtime and latency.
翻译:我们提出TabRepo,这是一个包含表格模型评估结果与预测的新数据集。该数据集收录了1206个模型在200个回归与分类数据集上的预测结果及评估指标。我们通过多种方式展示了该数据集的价值:首先,利用预计算的模型预测结果,可在无需额外计算成本的情况下进行集成学习,同时实现对超参数优化与当前AutoML系统的对比分析;其次,该数据集可直接用于迁移学习。特别地,我们证明了应用标准迁移学习技术后,能够在准确率、运行时间和延迟方面超越当前最先进的表格学习系统。