Large neural networks pretrained on web-scale corpora are central to modern machine learning. In this paradigm, the distribution of the large, heterogeneous pretraining data rarely matches that of the application domain. This work considers modifying the pretraining distribution in the case where one has a small sample of data reflecting the targeted test conditions. We propose an algorithm motivated by a recent formulation of this setting as an online, bilevel optimization problem. With scalability in mind, our algorithm prioritizes computing gradients at training points which are likely to most improve the loss on the targeted distribution. Empirically, we show that in some cases this approach is beneficial over existing strategies from the domain adaptation literature but may not succeed in other cases. We propose a simple test to evaluate when our approach can be expected to work well and point towards further research to address current limitations.
翻译:大规模神经网络在Web级语料库上的预训练是现代机器学习的核心。在此范式中,大规模异质性预训练数据的分布很少与应用领域的分布相匹配。本研究考虑在拥有反映目标测试条件的小样本数据情况下,对预训练分布进行修改。我们提出一种算法,该算法基于近期对该场景的在线双层优化问题建模。考虑到可扩展性,我们的算法优先计算那些最可能改善目标分布损失训练点的梯度。实验表明,在某些情况下,该方法优于领域适应文献中的现有策略,但在其他情况下可能不成功。我们提出一个简单测试来评估该方法预期有效的条件,并指出需进一步研究以解决当前局限性。