The scientific scale-up of large language models (LLMs) necessitates a comprehensive understanding of their scaling properties. However, the existing literature on the scaling properties only yields an incomplete answer: optimization loss decreases predictably as the model size increases, in line with established scaling law; yet no scaling law for task has been established and the task performances are far from predictable during scaling. Task performances typically show minor gains on small models until they improve dramatically once models exceed a size threshold, exemplifying the ``emergent abilities''. In this study, we discover that small models, although they exhibit minor performance, demonstrate critical and consistent task performance improvements that are not captured by conventional evaluation strategies due to insufficient measurement resolution. To measure such improvements, we introduce PassUntil, an evaluation strategy through massive sampling in the decoding phase. We conduct quantitative investigations into the scaling law of task performance. Firstly, a strict task scaling law is identified, enhancing the predictability of task performances. Remarkably, we are able to predict the performance of the 2.4B model on code generation with merely 0.05\% deviation before training starts. Secondly, underpinned by PassUntil, we observe concrete evidence of emergent abilities and ascertain that they are not in conflict with the continuity of performance improvement. Their semblance to break-through is that their scaling curve cannot be fitted by standard scaling law function. We then introduce a mathematical definition for the emergent abilities. Through the definition, we refute a prevalent ``multi-step reasoning hypothesis'' regarding the genesis of emergent abilities and propose a new hypothesis with a satisfying fit to the observed scaling curve.
翻译:大规模语言模型(LLMs)的科学扩展需要深入理解其缩放性质。然而,现有关于缩放性质的文献仅提供了不完整的答案:随着模型规模增大,优化损失呈现可预测的下降,符合既定的缩放定律;但任务层面尚无缩放定律,且任务性能在缩放过程中远不可预测。任务性能通常在小模型上仅显示微小提升,直到模型超过某个规模阈值后才会显著改善,这体现了“涌现能力”。本研究发现,尽管小模型表现微弱,但其任务性能存在关键且一致的提升,这些提升因传统评估策略的测量分辨率不足而未被捕捉。为测量此类提升,我们提出PassUntil,一种通过解码阶段大量采样的评估策略。我们对任务性能的缩放定律进行了定量研究。首先,我们识别出严格的任务缩放定律,增强了任务性能的可预测性。值得注意的是,我们能够在训练开始前以仅0.05%的偏差预测2.4B模型在代码生成任务上的表现。其次,基于PassUntil,我们观察到了涌现能力的具体证据,并确定它们与性能提升的连续性并不矛盾。它们看似突破的原因是缩放曲线无法用标准缩放定律函数拟合。我们随后为涌现能力给出了数学定义。通过该定义,我们反驳了关于涌现能力起源的流行“多步骤推理假说”,并提出了一个与观测缩放曲线高度吻合的新假说。