We propose a paradigm shift toward open-ended curriculum self-play: rather than learning to answer on a fixed prompt set, a unified policy learns to question: generating verifiable problems, solving them, and turning verifier feedback into self-improvement without human-annotated solutions. We introduce ANCORA, in which the policy alternates between a Proposer that synthesizes novel specifications and a Solver that produces verified solutions, anchored by three load-bearing mechanisms: a two-level group-relative update coupling Proposer advantages across specifications with Solver advantages across solution attempts; iterative self-distilled SFT projecting the base model onto its valid-output manifold before RL; and a UCB-guided Curriculum DAG whose policy-induced problem set can provably expand under self-composition. Without these stabilizers, sparse verifier feedback drives Proposer collapse even under MLRL-aligned rewards; with them, ANCORA bootstraps a verifiable curriculum from zero human solutions. Instantiated in Verus, ANCORA lifts Dafny2Verus pass@1 from a 26.6% SFT baseline to 81.5% in test-time training (TTT, 0-shot), outperforming PSV self-play by 15.8 points despite PSV's 1-shot inference; in a transfer setting, training from Dafny2Verus seeds yields 36.2% and 17.2% pass@1 on held-out MBPP and HumanEval.


翻译:我们提出了一种向开放式课程自我博弈的范式转变:统一策略不再在固定提示集上学习回答,而是学习提问——生成可验证的问题、求解这些问题,并将验证器反馈转化为无需人工标注解决方案的自我改进。我们引入ANCORA方法,其中策略交替执行两类角色:综合新规范的提案者(Proposer)与生成已验证解的求解者(Solver),并由三个承载机制锚定:双层级群组相对更新机制——将提案者跨规范的收益与求解者跨解尝试的收益耦合;迭代式自蒸馏SFT——在强化学习前将基模型投影至其有效输出流形;以及基于UCB引导的课程有向无环图(Curriculum DAG)——其策略生成的问题集可在自组合下得到可证明的扩展。缺少这些稳定机制时,即使采用MLRL对齐的奖励,稀疏的验证器反馈也会导致提案者崩溃;添加这些机制后,ANCORA能从零人工解决方案自举出可验证课程。在Verus框架中实例化后,ANCORA将Dafny2Verus的pass@1从26.6%的SFT基线提升至测试时训练(TTT,零样本)的81.5%,尽管PSV推理时采用一次样本而ANCORA为零样本,仍较PSV自我博弈提升15.8个百分点;在迁移设置中,基于Dafny2Verus种子训练在保留的MBPP和HumanEval数据集上分别取得36.2%和17.2%的pass@1。

0
下载
关闭预览

相关内容

自监督学习理论
专知会员服务
57+阅读 · 2022年8月23日
可解释强化学习,Explainable Reinforcement Learning: A Survey
专知会员服务
133+阅读 · 2020年5月14日
浅谈主动学习(Active Learning)
凡人机器学习
32+阅读 · 2020年6月18日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
23+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
17+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
国家自然科学基金
10+阅读 · 2012年12月31日
VIP会员
最新内容
对抗环境下超视距目标打击的情报支援
专知会员服务
3+阅读 · 今天14:49
《无人机对海面作战影响评估》
专知会员服务
11+阅读 · 7月21日
印度精确打击与指挥架构的断层
专知会员服务
6+阅读 · 7月20日
美空军AI完成F-16战斗机自主空战历史性试飞
专知会员服务
6+阅读 · 7月20日
相关VIP内容
自监督学习理论
专知会员服务
57+阅读 · 2022年8月23日
可解释强化学习,Explainable Reinforcement Learning: A Survey
专知会员服务
133+阅读 · 2020年5月14日
相关资讯
浅谈主动学习(Active Learning)
凡人机器学习
32+阅读 · 2020年6月18日
相关基金
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
23+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
17+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
国家自然科学基金
10+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员