LLM-driven software engineering agents have become a central testbed for real-world language-model capability, yet their training remains limited by the availability of high-quality SWE tasks. Existing synthetic data methods typically create tasks through fixed mutation or bug-injection procedures, making the resulting distributions largely independent of the agent's own weaknesses and training progress. We introduce Socratic-SWE, a closed-loop self-evolution framework that reuses the agent's historical solving traces as a source of training signal. Rather than treating traces only as evidence for reward computation, Socratic-SWE distills them into structured agent skills that summarize recurring failures and effective repair patterns. These skills then guide the generation of targeted repair tasks in real repositories. Candidate tasks are checked through execution-based validation and scored with a solver-gradient alignment reward, so that the retained tasks are both verifiable and useful for improving the Solver. The updated Solver produces new traces, enabling the task curriculum to adapt over successive rounds. Across SWE-bench Verified, SWE-bench Lite, SWE-bench Pro, and Terminal-Bench 2.0, Socratic-SWE consistently improves over self-evolving baselines under the same compute budget, reaching 50.40% on SWE-bench Verified after three iterations. These results suggest that solving traces can serve as a scalable substrate for self-evolving SWE agents.


翻译:基于大语言模型的软件工程智能体已成为评估真实世界语言模型能力的核心测试平台,但其训练仍受限于高质量SWE任务的可用性。现有合成数据方法通常通过固定变异或缺陷注入程序创建任务,导致生成的数据分布与智能体自身弱点及训练进程基本无关。我们提出Socratic-SWE,一种闭环自演化框架,通过复用智能体的历史求解轨迹作为训练信号源。不同于仅将轨迹作为奖励计算的证据,Socratic-SWE将其提炼为结构化智能体技能,用以总结重复性失败与有效修复模式。这些技能进而指导在真实代码仓库中生成针对性修复任务。候选任务通过基于执行的验证检查,并以求解器梯度对齐奖励进行评分,从而保留既可通过验证又有助于改进求解器的任务。更新后的求解器生成新轨迹,使任务课程能够在连续迭代中自适应调整。在SWE-bench Verified、SWE-bench Lite、SWE-bench Pro及Terminal-Bench 2.0基准测试中,Socratic-SWE在相同计算预算下持续优于自演化基线方法,经三次迭代后在SWE-bench Verified上达到50.40%。这些结果表明,求解轨迹可作为自演化SWE智能体的可扩展训练基板。

0
下载
关闭预览

相关内容

智能体,顾名思义,就是具有智能的实体,英文名是Agent。
综述 | Self-Evolving Coding Agents:自进化编程智能体
伯克利最新《智能体 AI (Agentic AI)》课程
专知会员服务
50+阅读 · 3月1日
智能体工程(Agent Engineering)
专知会员服务
40+阅读 · 2025年12月31日
浅谈群体智能——新一代AI的重要方向
中国科学院自动化研究所
44+阅读 · 2019年10月16日
【知识图谱】知识图谱+人工智能=新型网络信息体系
产业智能官
14+阅读 · 2018年11月18日
变分自编码器VAE:一步到位的聚类方案
PaperWeekly
25+阅读 · 2018年9月18日
NLP中自动生产文摘(auto text summarization)
机器学习研究会
14+阅读 · 2017年10月10日
群体智能:新一代人工智能的重要方向
走向智能论坛
12+阅读 · 2017年8月16日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
21+阅读 · 2013年12月31日
国家自然科学基金
19+阅读 · 2008年12月31日
VIP会员
最新内容
《国防与国际安全中的量子-人工智能融合》报告
专知会员服务
1+阅读 · 今天14:44
致命七类无人机:无人机时代的演进型合成兵种
《异构无人水面艇集群作战自主制导算法》130页
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
21+阅读 · 2013年12月31日
国家自然科学基金
19+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员