Larger models learn tasks smaller models do not. What drives this phenomenon? We develop a simple phenomenological argument that power-law scaling already suggests that a larger model will be able to learn a part of the data distribution that a smaller model fails to learn, even with infinite training data. To validate this claim and identify its causes, we study the effects of model scaling on a synthetic setup consisting of a mixture of tasks that show monotonic scaling curves. The results point to a data-induced competition over resources (neurons). Specifically, smaller models allocate their neurons to high frequency or low complexity tasks, and so they learn solutions that perform poorly on rare and complex tasks. Moreover, this happens even when solutions capable of expressing the desired task exist. We then assess how a larger model circumvents this data-centric bottleneck, finding that it traces to a reduced interference mechanism: larger models can allocate enough resources to common tasks that the gradient updates for those tasks become weak, which means that they do not overwrite rare-task features as they slowly accumulate. Finally, to further validate these claims, we pretrain OLMo models (4M to 4B parameters) on novel tasks of varying frequency and complexity. The results mirror those from our synthetic data experiments: only the larger OLMo models learn the infrequent and complex tasks, and these larger models embed more task features in their representations and show less gradient interference between tasks. Overall, we offer a data-centric account of why larger models learn tasks that smaller models fail to. This helps explain why larger models are better in practice, and it can inform practical questions concerning model sizing and training data mixtures.
翻译:更大模型能学到较小模型无法掌握的任务。是什么驱动了这一现象?我们提出一个简单的现象学论证:幂律缩放已表明,即便在无限训练数据下,更大模型也能学习到较小模型无法捕获的部分数据分布。为验证此论断并探究其成因,我们研究模型缩放对一组包含单调缩放曲线的混合任务合成实验的影响。结果指向数据驱动的资源(神经元)竞争:具体而言,较小模型会将神经元分配给高频或低复杂度任务,从而学到对稀有和复杂任务表现不佳的解决方案。此外,即使存在能表达所需任务的解,这一现象依然发生。我们进而评估更大模型如何规避这一数据瓶颈,发现其根源在于干扰机制减弱:更大模型能为常见任务分配足够资源,使这些任务的梯度更新变弱,从而在稀有任务特征缓慢累积时避免覆盖它们。最后,为进一步验证这些论断,我们在不同频率与复杂度的新颖任务上预训练OLMo模型(参数从4M到4B)。结果与合成数据实验相呼应:仅更大规模的OLMo模型能学到低频复杂任务,且这些更大模型在其表征中嵌入更多任务特征,并表现出更少的任务间梯度干扰。总体而言,我们提供了一个以数据为中心的视角,解释为何更大模型能学到较小模型无法掌握的任务。这有助于理解为什么更大模型在实践中表现更优,并为模型规模设计与训练数据配比的实践问题提供指导。