Large language models (LLMs) with enormous pre-training tokens and parameter amounts emerge abilities, including math reasoning, code generation, and instruction following. These abilities are further enhanced by supervised fine-tuning (SFT). The open-source community has studied on ad-hoc SFT for each ability, while proprietary LLMs are versatile for all abilities. It is important to investigate how to unlock them with multiple abilities via SFT. In this study, we specifically focus on the data composition between mathematical reasoning, code generation, and general human-aligning abilities during SFT. From a scaling perspective, we investigate the relationship between model abilities and various factors including data amounts, data composition ratio, model parameters, and SFT strategies. Our experiments reveal that different abilities exhibit different scaling patterns, and larger models generally show superior performance with the same amount of data. Mathematical reasoning and code generation improve as data amounts increase consistently, while the general ability is enhanced with about a thousand samples and improves slowly. We find data composition results in various abilities improvements with low data amounts, while conflicts of abilities with high data amounts. Our experiments further show that composition data amount impacts performance, while the influence of composition ratio is insignificant. Regarding the SFT strategies, we evaluate sequential learning multiple abilities are prone to catastrophic forgetting. Our proposed Dual-stage Mixed Fine-tuning (DMT) strategy learns specialized abilities first and then learns general abilities with a small amount of specialized data to prevent forgetting, offering a promising solution to learn multiple abilities with different scaling patterns.
翻译:具有海量预训练令牌和参数数量的大语言模型(LLM)涌现出多种能力,包括数学推理、代码生成和指令遵循。这些能力通过监督微调(SFT)得到进一步强化。开源社区已针对每种能力进行专门SFT研究,而专有LLM则具备全能型能力。研究如何通过SFT解锁多能力模型至关重要。本研究聚焦于SFT过程中数学推理、代码生成与通用人类对齐能力间的数据组成问题。从缩放视角出发,我们探究模型能力与数据量、数据组成比例、模型参数及SFT策略等多元因素的关系。实验表明,不同能力呈现不同缩放模式,同等数据量下大型模型表现更优。数学推理与代码生成能力随数据量增加持续提升,而通用能力在约千条样本后提升速度趋缓。我们发现低数据量时数据组成能促进多能力提升,但高数据量时能力间存在冲突。实验进一步表明组成数据量对性能有影响,而组成比例的作用不显著。针对SFT策略,评估发现顺序学习多种能力易导致灾难性遗忘。我们提出的双阶段混合微调(DMT)策略,先学习专项能力再辅以少量专项数据学习通用能力以防止遗忘,为学习具有不同缩放模式的多种能力提供了有效解决方案。