Progress in long-context reasoning for large language models (LLMs) has lagged behind other recent advances. This gap arises not only from the intrinsic difficulty of processing long texts, but also from the scarcity of reliable human annotations and programmatically verifiable reward signals. In this paper, we propose SPELL, a multi-role self-play reinforcement learning framework that enables scalable, label-free optimization for long-context reasoning. SPELL integrates three cyclical roles-questioner, responder, and verifier-within a single model to enable continual self-improvement. The questioner generates questions from raw documents paired with reference answers; the responder learns to solve these questions based on the documents; and the verifier evaluates semantic equivalence between the responder's output and the questioner's reference answer, producing reward signals to guide continual training. To stabilize training, we introduce an automated curriculum that gradually increases document length and a reward function that adapts question difficulty to the model's evolving capabilities. Extensive experiments on six long-context benchmarks show that SPELL consistently improves performance across diverse LLMs and outperforms equally sized models fine-tuned on large-scale annotated data. Notably, SPELL achieves an average 7.6-point gain in pass@8 on the strong reasoning model Qwen3-30B-A3B-Thinking, raising its performance ceiling and showing promise for scaling to even more capable models. Our code is available at https://github.com/Tongyi-Zhiwen/Qwen-Doc.


翻译:大型语言模型(LLM)在长上下文推理方面的进展滞后于其他近期突破。这一差距不仅源于处理长文本的内在困难,也源于可靠人工标注与可程序化验证奖励信号的稀缺性。本文提出SPELL——一种多角色自博弈强化学习框架,能够实现可扩展、无标注的长上下文推理优化。SPELL在单一模型中整合了提问者、应答者与验证者三个循环角色,以支持持续自我改进:提问者根据原始文档生成附带参考答案的问题;应答者学习基于文档解答这些问题;验证者则评估应答者输出与提问者参考答案之间的语义等价性,产生指导持续训练的奖励信号。为稳定训练过程,我们引入了自动课程学习机制以逐步增加文档长度,并设计了能根据模型动态能力调整问题难度的奖励函数。在六个长上下文基准测试上的大量实验表明,SPELL能持续提升不同LLM的性能,且优于基于大规模标注数据微调的同规模模型。值得注意的是,SPELL在强推理模型Qwen3-30B-A3B-Thinking上实现了平均7.6分的pass@8提升,突破了其性能上限,并展现出向更强大模型扩展的潜力。代码已开源:https://github.com/Tongyi-Zhiwen/Qwen-Doc。

0
下载
关闭预览

相关内容

强化学习增强的大型语言模型:综述
专知会员服务
53+阅读 · 2024年12月17日
基于模型的强化学习综述
专知
42+阅读 · 2022年7月13日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
VIP会员
最新内容
深入Project Maven:为何人工智能在战场上依然失灵
锻造未来士兵:外骨骼、基因工程与赛博格
专知会员服务
6+阅读 · 7月19日
《无人机蜂群通信技术研究》50页
专知会员服务
7+阅读 · 7月19日
战力倍增器:自主武器系统与乌克兰及加沙冲突
人工智能赋能战场情报:提速决策进程
专知会员服务
5+阅读 · 7月17日
《拥抱新兴技术:面向未来军官的教育革新》
专知会员服务
8+阅读 · 7月17日
相关基金
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员