Post-training pipelines that combine supervised fine-tuning (SFT) with reinforcement learning (RL) have emerged as the key recipe for transforming large language models (LLMs) into robust reasoners. We argue that this combined success is driven by compositional generalization, which we formalize through a hierarchical latent selection model. In this framework, reasoning traces are generated by a cascade of discrete latent selection variables corresponding to reusable atomic modules, including both skills (local operations) and routing mechanisms (how intermediate information is selected, reused, and composed). Within this model, we theoretically show that SFT and RL play asymmetric, complementary roles: SFT supplies the raw module materials in compositional traces, and RL decomposes those traces to identify the latent atomic modules and enable compositional generalization. We design controlled experiments to validate this theory. Our results demonstrate that RL can extract atomic modules from compound traces supplied by SFT and recombine them to solve new configurations. Moreover, we find that training on compound traces yields stronger generalization than training on isolated atomic modules. Finally, we investigate the relationship between SFT and RL data and identify an effective protocol in which SFT ensures coverage of all atomic modules through compositional traces, while RL focuses on novel compositions outside the SFT support to drive exploration.
翻译:结合监督微调(SFT)与强化学习(RL)的后训练流水线已成为将大型语言模型(LLMs)转化为稳健推理者的关键方法。我们认为,这种联合成功源于组合泛化,并通过分层潜在选择模型对其进行形式化。在该框架中,推理轨迹由一系列离散的潜在选择变量级联生成,这些变量对应于可重用的原子模块,包括技能(局部操作)和路由机制(中间信息的选择、重用与组合方式)。在此模型下,我们从理论上证明SFT与RL扮演着非对称的互补角色:SFT在组合轨迹中提供原始模块材料,而RL则分解这些轨迹以识别潜在原子模块并实现组合泛化。我们设计了受控实验来验证这一理论。结果表明,RL能从SFT提供的复合轨迹中提取原子模块,并将其重新组合以解决新的配置问题。此外,我们发现训练复合轨迹比训练孤立原子模块能带来更强的泛化能力。最后,我们探究了SFT与RL数据之间的关系,并提出了一种有效协议:SFT通过组合轨迹确保所有原子模块的覆盖,而RL则聚焦于SFT支持范围之外的新颖组合以驱动探索。