Pretraining a neural network on a large dataset is becoming a cornerstone in machine learning that is within the reach of only a few communities with large-resources. We aim at an ambitious goal of democratizing pretraining. Towards that goal, we train and release a single neural network that can predict high quality ImageNet parameters of other neural networks. By using predicted parameters for initialization we are able to boost training of diverse ImageNet models available in PyTorch. When transferred to other datasets, models initialized with predicted parameters also converge faster and reach competitive final performance.
翻译:在大规模数据集上预训练神经网络正成为机器学习领域的一个基石,但这一技术仅少数拥有大量资源的群体能够触及。我们致力于实现预训练民主化的宏伟目标。为此,我们训练并发布了一个单一的神经网络,该网络能够预测其他神经网络的高质量ImageNet参数。通过使用预测参数进行初始化,我们能够加速PyTorch中多种ImageNet模型的训练过程。当迁移到其他数据集时,以预测参数初始化的模型同样能更快收敛并达到具有竞争力的最终性能。