As language models scale up, it becomes increasingly expensive to verify research ideas because conclusions on small models do not trivially transfer to large ones. A possible solution is to establish a generic system that directly predicts some metrics for large models solely based on the results and hyperparameters from small models. Existing methods based on scaling laws require hyperparameter search on the largest models, which is impractical with limited resources. We address this issue by presenting our discoveries indicating that Maximal Update parametrization (muP) enables accurate fitting of scaling laws for hyperparameters close to common loss basins, without any search. Thus, different models can be directly compared on large scales with loss prediction even before the training starts. We propose a new paradigm as a first step towards reliable academic research for any model scale without heavy computation. Code will be publicly available shortly.
翻译:随着语言模型规模不断扩大,验证研究思路的成本日益高昂,因为小模型上的结论无法简单迁移至大模型。一种可行的解决方案是建立通用系统,仅基于小模型的实验超参数与结果,直接预测大模型的关键指标。现有基于标度律的方法需在最大规模模型上开展超参数搜索,这在资源有限时难以实现。本文通过揭示最大更新参数化(muP)能够在不进行任何搜索的情况下,对接近常见损失洼地的超参数实现标度律的精确拟合,从而解决了该问题。由此,不同模型可在大规模场景下直接比较,甚至在训练开始前即可通过损失预测进行评估。本研究提出一种新范式,旨在无需繁重计算的前提下,为任意规模模型的可靠学术研究迈出第一步。相关代码将很快公开。