We introduce Nemotron-Cascade 2, an open 30B MoE model with 3B activated parameters that delivers best-in-class reasoning and strong agentic capabilities. Despite its compact size, its mathematical and coding reasoning performance approaches that of frontier open models. It is the second open-weight LLM, after DeepSeekV3.2-Speciale-671B-A37B, to achieve Gold Medal-level performance in the 2025 International Mathematical Olympiad (IMO), the International Olympiad in Informatics (IOI), and the ICPC World Finals, demonstrating remarkably high intelligence density with 20x fewer parameters. In contrast to Nemotron-Cascade 1, the key technical advancements are as follows. After SFT on a meticulously curated dataset, we substantially expand Cascade RL to cover a much broader spectrum of reasoning and agentic domains. Furthermore, we introduce multi-domain on-policy distillation from the strongest intermediate teacher models for each domain throughout the Cascade RL process, allowing us to efficiently recover benchmark regressions and sustain strong performance gains along the way. We release the collection of model checkpoint and training data.
翻译:我们介绍了Nemotron-Cascade 2,这是一个开放的30B MoE模型(3B激活参数),具备顶尖的推理能力和强大的智能体能力。尽管其规模紧凑,其数学与代码推理性能已逼近前沿开放模型。该模型是继DeepSeekV3.2-Speciale-671B-A37B之后第二个在2025年国际数学奥林匹克竞赛(IMO)、国际信息学奥林匹克竞赛(IOI)及ICPC世界总决赛中达到金牌级表现的开放权重大语言模型,以20倍更少的参数实现了极高智能密度。相较于Nemotron-Cascade 1,关键技术进展如下:在精心整理的数据集上完成SFT后,我们大幅扩展了级联强化学习(Cascade RL)覆盖的推理与智能体领域范围。此外,我们引入了多域在线策略蒸馏技术,在整个级联强化学习过程中从各领域的最强中间教师模型获取知识,从而有效恢复基准性能退化并在过程中持续保持强劲性能提升。我们发布了模型检查点与训练数据的完整集合。