Competitive STEM examinations such as JEE and NEET require multi-step symbolic reasoning, precise numerical computation, and deep conceptual understanding across physics, chemistry, and mathematics. Recent large language models perform strongly on common reasoning benchmarks, yet they remain difficult to deploy at scale, where millions of student doubts demand domain-specific, consistently structured problem solving. We introduce Aryabhata 2, a reasoning-focused language model for competitive STEM examinations, trained via reinforcement-learning post-training. Using PhysicsWallah's internal question banks, we construct a high-quality training curriculum and post-train GPT-OSS-20B through reinforcement learning with verifiable rewards. Training combines prolonged reinforcement learning with broadened exploration via progressively larger rollout group sizes. We evaluate Aryabhata 2 on competitive examination benchmarks, including JEE Main, JEE Advanced, and NEET, as well as out-of-distribution reasoning datasets such as AIME, HMMT, MMLU-Pro, MMLU-Redux 2.0, and GPQA. Results show that Aryabhata 2 outperforms its base model GPT-OSS-20B on competitive STEM reasoning while requiring substantially fewer output tokens (up to 64\% fewer).
翻译:竞争性STEM考试(如JEE和NEET)需要多步符号推理、精确数值计算以及物理、化学和数学领域的深层概念理解。近期的大型语言模型在常见推理基准测试中表现强劲,但在大规模部署中仍面临挑战——当数百万学生提出疑问时,需要领域特定且结构一致的问题解决能力。我们提出专注于竞争性STEM考试推理的语言模型阿耶波多2,该模型通过强化学习后训练实现。利用PhysicsWallah的内部题库,我们构建高质量训练课程,并通过可验证奖励的强化学习对GPT-OSS-20B进行后训练。训练过程结合了长期强化学习与逐步增大rollout分组规模的广度探索。我们在竞争性考试基准(包括JEE Main、JEE Advanced和NEET)以及分布外推理数据集(如AIME、HMMT、MMLU-Pro、MMLU-Redux 2.0和GPQA)上评估阿耶波多2。结果表明,阿耶波多2在竞争性STEM推理任务中优于其基础模型GPT-OSS-20B,同时显著减少输出token数量(最高减少64%)。